Repeated submissions will eventually hand a language model a passing backtest. Qu, Chen and Wang put a frozen statistical referee between the agent and the book, then measure the price of doing so. The agent proposes investment factors but cannot alter the referee. Admission depends solely on rank-IC observed after each candidate arrives.

How factors earn admission

The referee evaluates a daily cross-sectional Spearman correlation. It ranks the factor at the close of t-1 and compares that ranking with day t returns across a point-in-time CSI 500 universe. A good factor delivers an average daily correlation of 0.02 to 0.03, with daily standard deviation near 0.12. Evidence accumulates slowly at that signal-to-noise ratio, often for years.

Before each outcome arrives, the referee stakes a fraction of the candidate's "capital" on that day's IC being positive. The stream first undergoes AR(1)-whitening. Its plug-in staking rule, aGRAPA, is capped at 0.8 of the no-bankruptcy bound. Capital cannot be expected to grow under a zero-edge null. The referee can therefore inspect it daily without the optional-stopping problem of a rolling t-stat.

Admission runs through online e-BH, the Benjamini-Hochberg rule applied to betting capital. The procedure uses a frozen universe of 2,000 slots at alpha 0.05. Capital must reach 40,000 for the first admission and 40,000/k for the k-th. Resubmitting a near-copy consumes another slot.

Retirement uses a second bet against the factor. A daily-restarted e-detector triggers when the evidence places the edge below a viability threshold of 0.015, with a run-length target of 1,260 days.

The authors pair three proposers, a round-robin script, a discounted-UCB bandit and an LLM controller, with the frozen referee and three leaky alternatives. Those alternatives are a peeking t-test inspected daily after 60 observations, an e-process with an age-relaxed bar, and a no-gate arm that admits every candidate after 126 days. Tests cover a synthetic world with planted truth, a probe-writing environment and 540 CSI 500 walk-forward cells. Campaigns begin in 2016-2019 and use data through 2026-08.

What does the guarantee cover?

The guarantee is narrower than the headline counts imply. The theorem controls the false-discovery rate under a zero conditional edge on the whitened stream. The paper defines "false admissions" as factors with realised mean IC below 0.015, which the authors acknowledge is a different target. Within the four-family library, every false admission from the frozen referee had an edge between zero and 0.015. Across its 100 cells, none admitted a factor with negative realised edge. The leaky referees admitted 10.8, 32.0 and 90.8 negative-edge factors per campaign under round-robin.

Two further limits come from the authors. For the marginal null under autocorrelation, the guarantee is measured rather than proved. Whitening produces null false acceptance of 1.8/2.0/2.4% at rho 0/0.2/0.4. The raw stream records 2.1/6.0/15.4%, against a nominal 5%. Real-data LLM arms are replays by models likely trained on the campaign years, placing them outside the proof's hypothesis.

The script admits 5-11 times fewer sub-threshold factors

With truth planted, the frozen referee admits 0.000 false factors per submission under all three proposers. Its leaky counterparts admit 0.26 to 0.85, with the same rate for the script, bandit and LLM. A hidden-retry attacker extracts +0.027 per submission from the peeking referee and nothing from the frozen one.

Results weaken on the CSI 500 four-family library. Under round-robin, the frozen referee makes 11.7 false admissions per campaign, while the leaky arms make 86.2, 105.2 and 196.0. Because the registered label overlaps the admission window, it favors referees that admit late. Restricting judgment to post-admission days raises the frozen count to 17.6. The resulting advantage is 5-11x under the script and 2.6-4.9x under the LLM. Against each family's own break-even, including 0.074 for reversal, the difference contracts to two- to threefold: 96.0 versus 180.4 to 276.2.

The proposer hardly changes false admissions under the frozen referee, which range from 9.4 to 13.3 per campaign across the bandit and three LLM families. Yield does change.

The LLM beats the script on yield in 6 of 6 family-by-library settings. It runs level with the bandit overall, while GPT trails the bandit in every nine-family start (2019: -0.082, Holm p 0.004). Any LLM advantage over the bandit occurs before 2020 across all three model families. Gemini records +0.026 before and -0.016 after. The authors interpret this as a prior eventually overtaken by data. They also argue that a model exposed to the period should improve most after 2020, whereas the sign reverses.

Memory contributes little. Six memory configurations, including none and shuffled, finish within 0.0025 yield of the full-memory arm, and 0 of 10 tests remain significant after Holm correction. Diagnostic probes are where the LLM adds something unavailable to a bandit. Authored probes reduce intervention regret in 3 of 6 evaluable model families, by 0.388, 0.248 and 0.233, with no significant difference in the other three.

Daily certification misses monthly momentum

At 15bp a side on the four-family library, the certified book posts net Sharpe of +0.33 to +0.50. The ungated book reaches +0.59 to +0.77. The abstract concedes the shortfall and attributes it to waiting. A true factor admitted by the frozen referee waits about 500 trading days, compared with 122-214 under the leaky referees. Earliness contributes the largest piece of the 2016 Sharpe decomposition, at +0.23 of the gap.

The paper's own counterfactual creates a harder problem. A fixed-horizon desk tests each candidate once at 500 days and applies BH to the same candidate sequence. Waiting time is essentially unchanged, 500 days versus 510. Yet net Sharpe rises by +0.12 to +0.24 in every start year. The fixed-horizon desk reaches +0.59 to +0.65, compared with +0.40 to +0.49 for the frozen book.

Its selections differ. The desk admits fewer reversal factors, 72.0 against 87.2, while allowing some momentum through. Momentum has daily IC of 0.001, rising to 0.009 at 63 days. During the 2016 round-robin runs, the ungated book carries 892 momentum sleeves earning +1.69bp a day. The frozen book carries six. A threshold rewarding the strongest daily statistic therefore selects reversal. Reversal's break-even is 0.074, and a reversal sleeve re-formed daily nets -19.6% a year.

The authors answer that the execution layer shelves 13% of certified admissions and 26% of ungated admissions, making Sharpe an unsuitable verdict on certification. Because the comparator tests the same one-day IC, it isolates the race to the anytime-valid threshold, a race won by the strongest daily statistic. Momentum is excluded at 0.001 daily IC. The paper accordingly concludes that certification should match the traded horizon. The authors also flag that this comparator is a counterfactual imposed on the frozen referee's candidate stream rather than a campaign of its own.

The result remains uninvestable, and the authors call the book "an instrument, not a strategy". Alpha t-stats remain below one in every certified four-family group. For the 2016 round-robin run, the certified book's net Sharpe becomes negative between 30 and 50bp a side. The test assumes a short position in a quintile of individual A-shares without a borrow list or fee. Its execution layer was built after an initial pass over the recorded cells, while the IC-to-return constant of 0.018 was estimated over the full panel.

Our US rebuild

We cannot trade CSI 500 here. Our version measures post-submission rank-IC on an annually point-in-time top-500 US universe selected by capitalization, using a strict append-only walk-forward. It exercises the same mechanism in a different market. Our run therefore leaves the paper unreplicated and untested, and none of its false-admission counts, waits or Sharpe ratios transfer to our results.

A historical run also cannot establish the guarantee's live-deployment interpretation, especially with an LLM proposer trained on the period. We can reproduce the scripted and bandit proposers cleanly. That evaluation is still running.