Most scheduled refits in Dutta's replay should have stayed on the sidelines. Of 528 challengers, only 114 passed. The remaining 414 could not beat the incumbent by the required 1e-4 of paired negative log-likelihood.
The replay follows a production-style forecaster for crypto perpetuals. Ten-band order-book depth is sampled on a 30-second UTC grid, producing a 60-snapshot by 40-feature input tensor. The target is the 300-second forward transaction-price return, classified as down, neutral or up beyond a fixed five-basis-point band. Unlabeled inputs refresh the serving normalizer moments at rate 0.01. As labels mature after 300 seconds, plain SGD updates the three-class head at 3e-4. During each weekly refit, the incumbent therefore absorbs another week of information.
Dutta calls the proposed release policy Shadow Before Swap. It is a champion-challenger gate designed around that moving incumbent. At each weekly boundary, the full serving state is deep-copied, including the representation, head, causal normalizer, delayed-label queue and online optimizer. The clone is warm-refit on the previous 28 calendar days, split into 21 for training and 7 for validation. It receives a fresh normalizer and a function-preserving reparameterization of the head, preventing the change in input statistics from changing predictions by itself. Both branches then shadow the same next week off the serving path, each with a separate queue. After the week's labels mature, promotion requires a paired mean NLL advantage above a deadband of 1e-4, roughly 0.009% of baseline weekly NLL.
Four complete recursive replays establish the comparison. Maintenance never full-refits. Calendar refits and deploys immediately. Blind waits through the same shadow week, then promotes every challenger regardless of its result. SBS enforces the gate. The data cover Binance USD-M and COIN-M perpetuals on eight underlyings with three seeds. The first out-of-time episode runs for 22 weeks from 4 Aug 2025 and contains 6,319,220 scored examples. The second runs for 26 weeks from 5 Jan 2026 and contains 6,789,702. Rules were fixed using a separate 26 weeks ending 3 Aug 2025, seed 0. We hold crypto price bars only. We do not hold Binance perpetual metadata or ten-band order-book depth, so nothing here was re-run on our side.
Across the pooled 48 weeks, SBS cuts NLL by 0.1472% relative to calendar, 0.0755% relative to blind and 0.0428% relative to maintenance. The absolute gains per forecast are 0.001567, 0.000803 and 0.000455.
How much comes from waiting?
Delay explains about half the calendar gap. Blind promotion incurs the same one-week lag, and the SBS-versus-blind difference is only 0.0755% beside the full 0.1472% calendar contrast. Dutta calls the smaller contrast the cleanest measure of authorization. That reading is right.
It is also the least stable contrast from week to week. SBS wins only 13 of 22 weeks in the first episode, where its advantage is 0.0423% and the four-week block interval is [0.0146, 0.0689]. Its worst week against blind is -0.0354% in the first episode and -0.0593% in the second. The corresponding episode averages are 0.0423% and 0.1036%. Among the 288 mature decisions in the 26-week episode, one-week trial gain has Spearman rho 0.524 with three-week forward value and 74.0% sign agreement. Dutta draws a restrained conclusion: a one-week trial filters the damaging tail rather than ranking the full field.
The maintenance comparison carries more weight than its size suggests. It addresses whether refusing every refit would be safer. The pooled advantage is 0.0428%, and 21.6% of proposals do add value. After selection among the three baselines, however, the simultaneous 95% lower endpoint falls to 0.0116%, about a quarter of the point estimate. Dutta uses 48 aggregate weeks as the inferential sample. The 13 million forecasts establish service-scale impact and nothing further.
The value of 0.000455 per forecast
The accumulation case holds because NLL is additive. A maintenance gain of this size becomes about 455 log-loss units per million predictions. With 2,880 grid steps a day across sixteen streams, the panel produces a million forecasts in roughly three weeks. Dutta's more usable translation fixes coverage at 5% in the 22-week episode. Against calendar, the difference corresponds to about 200 additional correct high-confidence decisions per million opportunities. Calendar is the weakest of the three comparators, and this calculation charges nothing for crossing a spread on a perpetual.
Dutta makes the economic argument in the same two parts. Section 5.2 says additive service-scale value and the matched-coverage diagnostic support practical significance, while monetary value depends on a specified downstream decision rule. The paper presents a deployment policy rather than a forecasting architecture. Turning the scores into profit is a separate estimand that requires prespecified trading costs, fills and market impact. Dutta also identifies equal asset and contract weighting as a deliberate choice. Capital- or risk-weighted results would need weights frozen before outcomes.
Deployed-state turnover falls 78.4%.
The deadband controls turnover far more clearly than NLL. Widening it from 0 to 5e-4 reduces promotions from 56 to 42 in the first episode and from 66 to 55 in the second. Across those four replays, the NLL effect barely changes: 26-week calendar records 0.1742%, 0.1738%, 0.1701% and 0.1705%. The threshold changes how often the model swaps. Its NLL effect stays nearly fixed.
A broken candidate meets the gate
The deliberately misspecified Temporal-CNN experiment is the paper's strongest result. The same refit recipe is applied to a 32-channel dilated residual network whose refits perform well in validation and badly on forward labels. Calendar replacement sends NLL to 12.6506, while blind promotion reaches 15.8667. Maintenance remains at 1.07117. SBS rejects 236 of 240 proposals and finishes at 1.07141, retaining 99.98% of the incumbent. Its cost against maintenance is -0.0221%, which Dutta describes as the price of shadowing a candidate generator with almost nothing deployable. Any desk would pay two basis points of relative log loss to avoid the twelvefold blowup from calendar replacement.
The headline rests on the second episode
All effects are larger in the second episode: 0.1738% versus 0.1157% against calendar, 0.1036% versus 0.0423% against blind, and 0.0470% versus 0.0379% against maintenance. The second episode contributes 26 of the pooled 48 weeks.
Dutta classifies this as a retrospective historical stability episode because the period had previously supported a maintenance audit. His defence has substance: the recursive SBS outcomes remained unopened when the protocol was registered. Even so, the larger half of the headline comes from data previously touched by the research process. The rule is described as development-selected and frozen rather than preregistered. Development used 26 UTC weeks from 3 February to 3 August 2025 with seed 0, under the same replay system and configuration, before analysis of the evaluation periods.
Breadth remains confined to one venue. The 20-asset 2024 panel (+0.0821%, +0.0458%, +0.0270%) and the cross-entropy challenger (+0.0757%, +0.0508%, +0.0195%) each test a single dimension. The Coinbase screen contains 55,367 trial examples, using downsampled snapshots and midpoint labels. Dutta describes these as breadth and objective transfer rather than generalization. The distinction is appropriate.
A prospective live run would change my view if it froze economic weights in advance and priced release costs, exactly the future work Dutta proposes. Until that evidence arrives, I would act on the governance half of the paper.