Size a position off the half-width of a conformal prediction interval and you get a risk gauge that stays honest out of sample and, on its own, does not make money once the search window ends.
The denominator is the whole trick
The construction is one substitution inside a textbook rule. For each of eight liquid US ETFs (SPY, QQQ, DIA, MDY, GLD, SLV, USO, DBC), an expanding-window ridge regression (lambda=10, refit every 21 days) maps four features (momentum over 21, 63 and 252 days, plus a 20-day EWMA volatility) to the forward 21-day return sum. The nonconformity score is the plain absolute residual. The 75% conformal quantile of the last 500 landed scores, geometrically shrunk toward an expanding anchor, becomes a scale: sigma = q_eff / 1.2816. Position size is fractional Kelly, f = 0.15 * mu / sigma^2, winsorised at plus or minus 0.75 per asset, then renormalised so gross exposure never exceeds 2.0. Wider interval, smaller bet. A calibrated width is an estimate of how wrong the forecast tends to be, and Kelly wants exactly that quantity in its denominator.
On the development window (2016-2021, 1,511 days, net of 5bps per unit turnover, no financing) the paper reports 28.45% annualised net log growth at Sharpe 1.336 with a 27.68% drawdown for the plain configuration, and 25.84% at Sharpe 1.386 with a 20.26% drawdown once a drawdown dial is added. Realized coverage was 0.7483 against 0.7500 nominal.
Why did the slow width win?
The result that earns the paper its keep is a negative one against the conformal-for-time-series literature. Every device that makes the interval adapt faster lost development growth: volatility-scaled scores -3.3pp, adaptive conformal inference -1.6pp, recency-weighted calibration -1.4pp, asymmetric CQR -1.6pp, pooled -2.5pp, Mondrian -2.4pp. A sigma frozen on pre-2016 residuals for six years (0.2669) still beat a rolling residual standard deviation (0.2476). At matched gross the conformal quantile beat the plug-in standard deviation by 2.1 percentage points a year and 0.08 of Sharpe, with mean absolute deviation sitting between them.
The explanation is mechanical and I find it convincing. A sizing map integrates the interval, so it is charged for the estimation variance of q, not rewarded for local sharpness. The growth ordering across estimators is their tail-sensitivity ordering: the standard deviation's influence is quadratic in the residual, the MAD's linear, the quantile's bounded. Understating fat-tail sigma keeps more weight on the profitable assets, which pays only because the next fact holds.
A cap doing the heavy lifting
Pre-cap gross leverage averages 4.28 and the 2.0 cap binds on 97.7% of development days. So kappa is inert, correlation is inert, and constant de-levering changes nothing. The book sits pinned at maximum gross, on the rising part of the growth curve, essentially always. A drift-only forecast (pure intercept, zero features) already reaches 82% of the final result, and the four equity ETFs supply +0.237 of Config A's +0.312 total arithmetic contribution. This is substantially a levered long-equity book. A post-hoc four-equity 2x bar beat Config A on development growth (0.2934 versus 0.2845), though at a 60.9% drawdown against 27.68%. The winsorisation clip is load-bearing too: renormalising before clipping ("water-filling") costs 4.0pp.
What we ran, and what it showed
Two things about our build before any of our figures. We could not replicate the exact experiment: the paper uses a frozen Kaggle price snapshot starting in 2006, while we ran on daily ETF close history that generally begins around 2010, so exact numbers were never on the table. And we did not reproduce the autonomous agent search or the sealed-lockbox protocol, which are not needed to trade the disclosed configuration; we evaluated the fixed Config B rule directly.
Our backtest covers 2020-01-01 to 2025-10-08 and compounds at about 6.7% CAGR with a Sharpe of 0.41 and a 40% maximum drawdown. Those are our numbers, from our one automated pass. Set against the paper's own figures, they are nowhere near the development Config B headline (25.84% growth, Sharpe 1.386, 20.26% drawdown) and land almost exactly on the paper's own sealed 2022-2024 lockbox for Config B (7.01% growth, Sharpe 0.422). Our window is dominated by the same 2022-2024 regime the authors sealed away, where they too reported growth failed to transfer, so the overlap is roughly what you would expect. Two open contradictions in our setup likely widen the gap further. The warmup arithmetic (750 days plus 21-day label landing plus 500 landed scores implies roughly 1,271 days before a fully formed scale exists) means much of our window trades on immature sigma if it was seeded only from 2020. And a stray 0.125 position cap in our framework metadata, if it bound instead of the paper's 0.75 winsorisation, would have kept us well below the cap-saturated regime the authors say drives everything. Our 40% drawdown, deeper than the paper's lockbox 31.7%, is the piece I can pin least: one candidate is that our equity curve straddles both the 2020 and 2022 selloffs while the dial is documented to sell the bottom during rebounds. We cannot fully close the gap from what we can see, and a thin result on a different window is evidence about our implementation first.
The dial as a risk gauge
Using the trailing one-sided downside miscoverage rate to raise growth failed across 40-plus configurations, because the signal peaks after a shock and sells the rebound. As a leverage dial it instead cut the development drawdown from 27.68% to 20.26% while lifting Sharpe to 1.386, beating all 40 circular-shift placebos on drawdown (p=0.024), and a constant-leverage control at matched gross left the drawdown unchanged at 27.68%, so the effect is timing rather than de-levering. Out of sample the calibration transferred (0.745 against 0.750) and the economics did not: both configurations ranked last of eleven entries on Sharpe and Calmar, and at a 4% financing charge Config A's lockbox growth falls to roughly +4.7% a year, below every unlevered passive bar. The width is a clean, slow, per-asset noise estimate. Turning it into wealth under a binding cap is the part that did not survive contact with 2022.
Worth reading for the eight refuted hypotheses and the reason slow scale beats fast scale in a sizing map. Not worth trading as printed, which the author says plainly.