Changing between two fixed AI recommendations moves the volatility of hypothetical pension portfolios by 2.146 points. Roughly 37% of the distance between those recommendations survives in the allocations people finally choose. Choi, Kim, Kovach, Lee, Shin and Tzavellas estimate a 95% interval around that fraction. It is the paper's most useful result.

We ran no backtest here, and could not. The paper studies allocations from a Korean workplace pension menu, rather than a tradable strategy. Reproduction would require a controlled behavioral experiment using generated recommendations alongside participant-level initial and revised allocations. Substituting a US ETF menu with similar risk grades could recreate the axis arithmetic. It would capture none of the behavior that makes the measure useful.

The randomized advice

Participants choose from a workplace defined-contribution pension menu containing eleven products from Ha et al. (2019), constructed to resemble actual Korean retirement products. Each row reports gross returns over three months, six months, one year, two years and three years. It also gives a risk grade from 1 to 5 and a product standard deviation ranging from 0.09 for the sole principal-guaranteed product to 28.24 for the riskiest. The researchers derive expected returns by geometrically annualizing the single three-year observation. They treat the products as uncorrelated because the menu supplies no correlations. Their figures use returns net of 1.2% expected inflation and a 0.6% fee, while the menu displays gross returns.

The experiment covers 400 employed South Koreans aged 35 to 55, all enrolled in DC plans. Each participant divides a hypothetical balance among the eleven products using integer weights that sum to 100. They then receive an AI-generated recommendation and can revise the portfolio. Data collection took place in Spring 2024, on desktop only, with 100 participants per cell and gender-by-age quotas. The study was preregistered on AsPredicted. A retirement-wealth simulation determined a bonus of 500 to 30,000 KRW, making volatility relevant to the payoff.

Randomization has two dimensions. The content treatment assigns either an aggressive vector, (10,10,5,15,0,0,10,20,0,0,30), generated by GPT-4 Turbo, or a conservative vector, (20,20,10,10,10,10,5,5,0,5,5), generated by GPT-4o. Between the vectors lies a 3.123 percentage-point expected-return gap and a 6.256-point volatility gap. The aggressive portfolio's Sharpe ratio is 0.071 lower. The second dimension determines whether the figures arrive with a short verbal rationale.

The measurement earns its keep. Every portfolio is projected onto the line connecting the recommendations, with conservative normalized to 0 and aggressive to 1. The randomized difference in mean final position therefore gives the fraction of the advice gap that enters participants' portfolios.

Pass-through stays near 0.368

Headline pass-through is 0.368, 95% CI [0.303, 0.433], n=400. Both zero and one are rejected at p<0.001. Alternative specifications change little: the change-score estimate is 0.372, while demographic, comprehension and preference controls produce 0.364. Pooling at product level and clustering by participant gives a slope of 0.372.

Transmission also looks fairly even across the risk measures. Assignment to the aggressive recommendation adds 1.221 points of expected return (SE 0.133) and 2.146 points of standard deviation (SE 0.266). Those effects represent 39.1% and 34.3% of the corresponding gaps. The principal-protected share declines 3.521 points and the low-risk share declines 7.739. High-risk rises 13.744, while very-high-risk rises 8.083. Pass-through across these categories spans 34.4% to 40.4%. Breadth responds less: participants hold 0.836 fewer products against a three-product gap, or 27.9% transmission.

Two outcomes show no detectable transmission. The recommendations differ by 0.055 in HHI and by 10 points in maximum product share, yet the causal effects are 0.003 (SE 0.011) and 0.701 points (SE 1.163). Risk-adjusted performance is similarly flat. Against a Sharpe gap of -0.071, the estimated effect is +0.012 (SE 0.012). Implied pass-through is -0.173 with a standard error of 0.175. Participants follow the risk setting while largely preserving portfolio structure.

The authors themselves flag the limits of the Sharpe null. Their power calculations place the minimum detectable Sharpe effect at about 0.06 to 0.08, leaving a modest genuine improvement below the study's reach. Sharpe also depends on the zero-correlation assumption. Setting a common pairwise correlation of 0.2 flips the ranking of the recommendation vectors. It also raises the share of participants whose baseline objectively dominates the advice from 3% to 6%. The Sharpe-gap sign should be read as an artifact of the correlation assumption. The usable finding is the absence of a large effect within a detection floor of 0.06 to 0.08.

How far people move

Of 400 participants, 323 revise their allocations. Among those revisers, 95% move toward the recommendation they saw, while 12 move away and three move orthogonally. Mean weight on the advice is 0.51, with a median of 0.52. A typical reviser therefore executes about half the proposed move. Directedness averages 0.65 and has a median of 0.76. Assigning the 77 non-revisers a value of zero lowers the full-sample figure to 0.41.

The placebo gives the result much of its credibility. The authors stack 4,400 asset-level revisions, then regress each revision on two gaps: one toward the recommendation received and another toward the unseen recommendation. The received gap has a coefficient of 0.369 (SE 0.025). The unreceived gap has a coefficient of 0.001 (SE 0.017, p=0.94). Their difference is significant at p<0.001, and the regression accounts for 29% of cell-level variation.

This is a demanding placebo because directions from baseline toward the two recommendations have a mean cosine of 0.70. The gap vectors are mechanically correlated, which makes the zero coefficient informative. Distance from the received recommendation drops 0.11, from 0.52 to 0.41, a 21% reduction at p<0.001. Distance toward the unassigned recommendation falls only 0.032.

The two-vector problem

Adding a short rationale produced no detectable change in pass-through. Its point estimate moved 0.089 in the other direction, p=0.179. Pass-through measured 0.324 without a rationale and 0.412 with one. The average rationale effect on axis position was -0.015 (p=0.653).

The authors end with a deployment claim: organizations may steer users toward selected risk profiles through their choice of model. Their design section acknowledges the limitation directly. The contrast "identifies the effect of assignment to these two specific recommendation portfolios; it does not identify a general effect of one GPT model relative to another."

That defence works on its stated terms. Whatever generated the vectors, the content contrast exists. The placebo shows that participants tracked the vector placed before them: 0.369 on the received gap, compared with 0.001 (SE 0.017) on the unreceived gap. Sampling uncertainty on the user side is explicit at 0.368, 95% CI [0.303, 0.433], n=400.

The model side consists of one draw each. Each recommendation is the modal allocation from 100 queries at temperature 0.5. The mode appeared 32 times for GPT-4 Turbo and 16 times for GPT-4o. The authors report the dispersion and say they do not test determinism. A plan sponsor ultimately cares about the product of the generator gap and pass-through. Here, only the second factor carries a standard error worth reading.

Perceived superiority also matters. Some 27% of participants believed their portfolios beat the AI on both return and risk. In fact, 3% did: 6% against the aggressive vector and 0% against the conservative one.

That belief reduces adherence. Within the full sample of 400, dominance-overconfident participants closed 0.049 less of the distance gap (p<0.01), or 0.036 less after full controls (p<0.05). Their revision magnitude was statistically identical, at 0.011 with p=0.73. Among the 323 revisers, directedness was 0.108 lower (SE 0.042), averaging 0.54 versus 0.69 for everyone else.

Baseline inefficiency changes a different margin. Participants whose baseline Sharpe fell below that of the received recommendation, 59% of the sample, revised 0.045 more. Mean revision magnitude is 0.28, making the coefficient roughly 16% of that amount. Among revisers, Sharpe-dominated participants assigned 0.131 more weight to the advice (SE 0.040). Their final distance was unchanged: the difference in gap closed was 0.001 (SE 0.015). The authors describe all these relationships as descriptive because beliefs were elicited after revision and reverse causality remains possible.

I would interpret the volatility result differently if the authors drew twenty recommendation vectors from each model and reported the distribution of return and volatility gaps, instead of relying on one mode apiece.