A drift slope of 0.96 means GPT-5.2 effectively leaves existing holdings alone.
The baseline simulation regresses the actual change in recommended equity share on passive drift, with a standard error of 0.01. A coefficient near one says the advice scarcely rebalances the portfolio already in place. Calvet, Campbell and Sodini estimate 0.5 between actual and passive portfolio changes in observational household data. Choukhmane, de Silva, Lin and Akuzawa interpret that comparison as roughly twice as much inertia in LLM recommendations. Against the paper's calibrated model, the printed slope is 0.03 (0.01). The baseline figure describes the two measures as essentially uncorrelated in the model. GPT-5.2 allocates new savings while leaving the accumulated stock of wealth where returns put it.
We reproduced none of these results. Replication requires the same LLM versions used by the authors, followed by the same two-stage prompt-to-advice-to-allocation process. Their calibrated lifecycle simulator was also unavailable in the market data we could use. We could only partly proxy the paper's household asset classes with US equity, bond and commodity ETFs, US stocks and crypto. That substitution would omit fixed-income and collectibles holdings, along with the entire household context. We therefore accepted no backtest, and all figures below belong to the authors.
The experiment behind the 0.96 deserves attention. The authors surveyed 1,000 U.S. adults through Prolific, asking each person to write three free-text prompts for an LLM: describe your financial situation, ask how much to spend, ask how to invest. After excluding 33 responses that failed Prolific's authenticity screen and 15 that failed a specificity check, 952 remain. The survey also records age, income, employment status, wealth, Big Five financial literacy and prior AI use. All three prompts together average 96.7 words, and 76.3% contain a number.
The authors place those prompts inside an annual lifecycle model in the Gourinchas-Parker and Cocco-Gomes-Maenhout tradition, closest to Choukhmane and de Silva's earlier work. SIPP supplies income dynamics and employment transitions. Mortality comes from the 2015 SSA tables. The equity process uses CRSP's total US market return series, specifically the value-weighted index from 1925 to 2006, while individual-stock returns use Bessembinder's 2018 moments. Taxes and Social Security follow 2025 rules. The calibration assumes a 6.4% equity premium, 20% log return volatility and a 2% real risk-free rate.
Four assets enter the model: bond, diversified equity, individual stock, other risky. Each year, every one of 1,000 simulated agents aged 22 to 89 receives a prompt drawn from one of 12 employment-by-age-by-income buckets. The text is updated with the agent's own state variables. GPT-5.2 responds at reasoning effort Low under a 200-word cap. GPT-5 Mini converts that prose into dollar consumption, contributions, withdrawals and transfers consistent with the budget identity. Gemini 3 Flash and GPT-5.6 Terra serve as checks.
A stronger household default
One-third of respondents report zero equities, and their observed equity share averages around 30%. Following the advice pushes participation close to universal and produces recommended equity shares averaging about 65%. The share begins declining after roughly age 45. Risky holdings concentrate in diversified funds: by retirement, non-diversified participation reaches about 35% for individual stocks and 50% for other risky assets, yet conditional positions amount to only 2 to 3% of financial wealth.
More than 20% of respondents in every age bracket hold under $10,000. Virtually every simulated agent following the advice passes $10,000 by age 30. Among employed agents aged 60 to 64, mean liquid wealth reaches $2,147,938.
Prompt content also moves the recommendations in the direction theory predicts. Mentioning macroeconomic conditions raises the recommended saving rate about 4pp and lowers equity share about 2pp. Mentioning financial hardship lowers the saving rate about 10pp.
State dependence exposes the weakness
SMM fitted to 324 wealth-to-income and equity-share percentile moments returns risk aversion of 5.3, an ordinary estimate. The discount factor is 1.034, an implausible one. Across ages 22 to 64, employed agents under the advice have a simulated net saving rate of 17%, compared with -7% in the optimizing model. The authors discuss bequests and paternalism before giving the candid interpretation: taken literally, the advice can be rationalized only through an unrealistically high discount factor. Their SMM has two free parameters and uses the identity weighting matrix, leaving some of the 1.034 attributable to misspecification.
Individual recommendations lean heavily on heuristics. Multiples of 10% account for 31.0% of recommended saving rates, while 34% of saving amounts are multiples of $5,000. More than 98% of recommended retirement withdrawals land at or below 4% of assets. Following a simulated job loss, income falls roughly 50% and recommended consumption drops by about the same amount, despite the agent's large liquid balance.
A structured prompt written by the researchers fixes part of the problem. It supplies every state variable, states the return assumptions and directs the model to act as a professional adviser in the user's best interest. Round saving rates fall to 14.6%, and the share of withdrawals below 4% drops to 8.8%. The fitted discount factor moves to 0.990, with risk aversion at 4.7. Consumption also falls less after unemployment and reverses sooner.
Inertia remains. Under both prompts, the drift slope stays at 0.96 (0.01). Asking directly for numbers, while skipping the two-step process, leaves drift at 0.95 even as the heuristics weaken from 31.0% to 18.7% and from 98.3% to 92.4%. Gemini 3 Flash and GPT-5.6 Terra produce 0.97 and 0.98.
Across the academic prompt, the direct process and the two alternative models, inertia holds at 0.96, 0.95, 0.97 and 0.98. The optimizing model remains at 0.03.
How 1.50 points divide
Gender is absent from 83% of prompts. For those prompts, the authors randomly add either "I am a man" or "I am a woman", then regress the recommended change in diversified equity share on author gender and randomized label with bucket fixed effects. Shifting from a prompt written by a man and labeled male to one written by a woman and labeled female reduces recommended diversified equity share by 1.50pp in that period (0.21).
Demand, defined as what women wrote, accounts for -0.96pp (0.15). Supply, the randomized label alone, accounts for -0.54pp (0.15). The -0.19pp gap in non-diversified assets comes entirely from demand, with a label estimate of 0.01 and s.e. 0.03. The net saving rate gap is -2.63pp (2.34), statistically indistinguishable from zero. The paper explicitly rules out offsetting forces: both components are negative, with demand at -0.88pp and supply at -1.76pp, and neither is statistically distinguishable from zero.
For race, the supply effect is small and statistically insignificant, though less precisely estimated than the gender benchmark. Black versus White gives +0.69pp in total and a 0.03 label effect. Asian versus White gives -0.81pp in total, with -0.16 (0.18) from the label. A null response to an explicit label still leaves room for reactions to subtler markers such as names, dialect or contextual cues.
This decomposition is the paper's advance over existing label-audit work, including the 1.8pp gap across 33 LLMs reported by Foltyn and Olsson and cited here. Prompt content explains two-thirds of the combined gender effect, with the remainder coming from the model's response to a stated gender. Those demand-side differences persist as models improve, while model design can magnify or reduce them.
The effects compound. Simulations using women's prompts produce wealth at 60 that averages $60,000, or 4.1%, below the result from men's prompts. The s.e. on the log gap is 1.61%. A 2.94pp lower average equity share drives the difference, while the future value of cumulative net saving is nearly identical for the two groups. The level estimate has a standard error of 31,185, making the log result the firmer figure.
The authors attach two caveats. The -1.50pp estimate is a single-period draw rather than a decomposition of the lifetime gap. And in one-shot advice, running the full sequence five times for the same prompt produces a median within-person standard deviation of 6.5pp in recommended equity share and $3,220 in consumption. The translation step alone yields $335 and 1.6pp. The regression coefficient is precise; a given recommendation can vary widely.
Inertia is the result that endures. The paper's proposed mechanism is the LLM's focus on allocating new savings, with few recommendations to rebalance existing holdings. Supplying every state variable, including the current portfolio, does not dislodge it. The authors state the central limitation in their conclusion: the results describe what happens when people follow LLM advice, rather than what happens when they merely receive it. Whether people act on these recommendations, and how the effects compare with human advisors, robo-advisors and other guidance, remains open. They cite early evidence from Moss, Wegner and Zechman, where investors follow an AI adviser's recommendations and heavier user intervention accompanies worse risk-adjusted performance.