An LLM system prompt belongs in the model-risk controls of any desk using 10-Ks for research. Asaad, Mohamed, Zhang and Abdelsalam quantify why. DeepSeek-V3.1 read a Coca-Cola 10-Q from October 2023 and assigned +0.60 under a long-only role, then -0.60 under a short-seller role. The filing never changed. The resulting 1.20 spread spans much of a scale running from -1 to +1, even though the prompt explicitly defined that field as a neutral, user-invariant assessment of the evidence.

The audit design

Each call produces a two-score schema. The EvidenceScore is supposed to remain independent of persona or user, while the PersonalizedScore may adjust to either. Both lie in [-1,+1]. The output also includes a BUY/HOLD/SELL recommendation, rationales, a confidence figure and a risk score.

The audit has one question: with the filing held fixed, does the EvidenceScore change when only the user context changes?

Three conditions separate retrieval effects from prompt effects. Persona retrieval lets the investor role shape the retrieval query and the prompt, meaning a long-only investor and short seller may receive different chunks from the same document. Neutral retrieval runs one neutral query for each filing, then gives every persona the identical retrieved chunks. In the memory condition, the assistant becomes a neutral analyst and receives the investor mindset third-person through a user profile. The user might be described as tail-risk-focused and inclined to weigh covenants, leverage and adverse disclosures heavily, instead of receiving the instruction "You are an activist short-seller".

The corpus contains 3,575 SEC 10-K and 10-Q filings from S&P 500 constituents. Filing years run from 2013 to 2024, with the sample stratified across 11 GICS sectors and capped at six filings per firm. The authors retain MD&A, Risk Factors, and Quantitative and Qualitative Disclosures, then use top-k=8 retrieval within each document. Twelve models from five families run with deterministic decoding. Filing-fixed-effects regressions estimate the effects against the neutral frame, with clustering at the filing level and a 1,000-replication bootstrap.

Identical evidence, much of the effect remains

Retrieval looks like the obvious culprit until the retention results arrive. The pinning check works: under neutral retrieval, the median per-filing Jaccard overlap of retrieved chunk IDs across personas is exactly 1.000. Under persona retrieval, it is roughly 0.66.

Yet across the four headline matched pairs, the twelve-model panel retains 69% of the persona-retrieval effect on average. Median retention is 68%, with a range of 44% to 89%. The retained effect is substantial. Across the panel, the mean absolute persona effect on the evidence score reaches 0.20 and the median is 0.21. It ranges from 0.02 for gemma-4-31B-it to 0.40 for Llama-3.1-8B-Instruct.

The short-seller frame is particularly sticky. Among 8 of 12 models, retention is at least 74%, and the panel median reaches 89.1%. Across nine personas, the mean per-filing spread falls from 0.594 to 0.443. Fixing retrieval therefore produces a 25% reduction.

Rationales show the same ordering. For DeepSeek-V3.1, median cosine similarity with the neutral analyst's rationale on the same filing rises from 0.721 under persona retrieval to 0.805 under neutral retrieval, then 0.849 in the memory condition. The first two arms contain 32,175 rationale pairs; the memory arm contains 17,875. Their interquartile range contracts from 0.19 to 0.13. Although the embedding model is outside the twelve being audited, it recovers the same pattern.

Persona framing still changes retrieval, though less than the prompt-side effect. On DeepSeek-V3.1, the risk manager retrieves Risk Factors 22.3% of the time, compared with 13.9% for the neutral baseline.

Where the mindset appears matters

Recasting the identical mindset from assistant identity into a third-person user profile reduces the short-seller evidence-side coefficient by 51% to 87%. The panel median reduction is 73%. On DeepSeek-V3.1, the median across all five matched pairs reaches 80%.

A 2x2 ablation tests whether the reframing works because of the dual-output schema. On the two models included, profile-framing spillover remains at 16% to 39% of role-framing under both output formats.

The Coca-Cola filing makes the distinction concrete. Its neutral-retrieval spread remains 0.50, equal to 42% of the persona-retrieval spread. Under the memory condition, all three profiles shown in the paper fall within [+0.10,+0.20]. The bearish interpretation shifts into the personalized field, where the design intended it to appear.

All twelve models fall below the 0.5 no-routing benchmark for leakage. The panel mean is 0.19 and the median 0.18. Results range from 0.05 for Llama-3.3-70B-Instruct to 0.37 for Qwen3.6-35B-A3B, while gpt-oss-120b records 0.34. Leakage follows the checkpoint rather than parameter count. As the authors state, larger models are not uniformly safer, and MoE models are not uniformly more invariant.

The headline metrics omit the momentum pair because its base effect is near zero. Appendix counter-examples to the framing claim then concentrate among those weak-signal personas. The authors disclose both results without presenting the difference as a problem. Framing reduces bias in 86.5% of 37 meaningful cells and 68.3% of all 60. For an arbitrary persona, the second denominator gives the more honest success rate.

The return table is desk-unready

A quintile sort on EvidenceScore, held for 252 trading days, produces a panel-mean Q5 minus Q1 abnormal return of +2.96% against the S&P 500. The calculation uses roughly 3,393 filings. Persona results range from Value at +5.63% and LongOnly at +3.84% down to Momentum at +1.70% and RiskManager at +0.08%. Qwen3.5-35B-A3B leads the model results at +6.35%. DeepSeek-V3.1, the largest model in the panel, lands mid-pack at +3.12%.

Those returns are gross of everything. They include no transaction costs, borrow or implementation lag, and measurement begins on the filing date, assuming the text is available and tradable that same day. The universe includes firms that belonged to the S&P 500 at some point from 2013 to 2024. Because that membership criterion is resolved with hindsight, I read it as an upward bias on the long leg.

The exercise is a raw quintile sort without size, value, momentum or beta controls. We did not find t-statistics for the spreads. Its 252-day windows overlap heavily across firms and years. The authors explicitly say this analysis is not the paper's target; their point is that the score contains information while persona conditioning shifts its level. Treated as a strategy, it is a large-cap post-filing drift sort with hindsight membership and no costs.

Memorization decides whether the score means anything out of sample. Models released after most of the sample evaluate public 2013 filings that, in the authors' words, may have appeared in pretraining. An entity-masking control leaves the persona effect essentially unchanged, with retention at or above 97%. That protects the bias result, though the control covers two of the twelve models. It provides no evidence on whether the return spread reflects forecast or recall.

We hold no historical SEC filings corpus split into MD&A, Risk Factors and the quantitative disclosures, so we could not rerun any of this.

What holds up

The invariance finding survives. It uses a neutral reference instead of externally validated truth, allowing the paper to show that the score moved without identifying which reading was right. The authors acknowledge that limit. Movement remains the relevant object: two analysts can query the same system with different stored profiles, receive different "neutral" readings of the same 10-K, and never see the other's result.

Two desk-level details sharpen the concern. On DeepSeek-V3.1, residual spillover is 2 to 4 times larger in the sample's smallest bucket, S&P 500 names under $10B with n=695, than in the over-$50B bucket with n=791. The risk-manager frame is also the only persona with an inverted return spread, at -1.40% on DeepSeek and +0.08% on the panel. Yet it carries the strongest forward-volatility signal, z=-8.6 with a coefficient of -0.45. Under that role, the model reads risk well and gets direction backwards.

Before an LLM reading enters a portfolio discussion, check retention under fixed evidence and leakage into the personalized field for each model. Parameter count predicts neither. The paper supplies both metrics and the regressions needed to calculate them.

I would take the return column seriously after the quintile sort is rerun on a point-in-time universe, net of costs, using a model whose pretraining cutoff predates the filing.