A desk should demand the corrected deflation benchmark before trusting any backtest. With the fixed fallback variance of 1 that the authors had been shipping, a strategy chosen from 100 trials needed an annualized Sharpe of roughly 3 to pass the Deflated Sharpe gate. Substitute Lo's null sampling variance of the Sharpe estimator, 1 over years, and the hurdle falls to about 1.9 on five years of history and about 1.3 on ten. The true-positive rate for genuine synthetic strategies climbed from 1% to 41%. Pure noise still produced a false-positive rate of exactly zero. If every strategy you ran received a deflated Sharpe of zero, the benchmark caused it.

Five gates, one display

MinervaScore grades the strategy already chosen by the parameter search. It stays outside that search loop by design: "a validation metric is informative only if it is evaluated on a strategy that was not selected to maximize it." A logical AND across five gates determines the binary Robustness Seal:

The abstract gives the authors' own comparison: "Its improvement over the GT-Score proxy and the gates-passed baseline is modest, and it remains close to the corrected DSR-alone baseline."

Each gate is converted into a signed distance from its threshold and standardized using a frozen cross-sectional dispersion. A weighted inverse-normal blend combines those margins. The result then passes through a normal CDF with a conservative offset of 0.5.

The verdict controls the display mapping. Sealed strategies map linearly to 80 to 100. Failing strategies receive 80 times their percentile in a frozen reference population, capped one tenth below 80, leaving 79 as the highest failing display. Four easy passes cannot lift the display above 80 if one gate fails. The invariant recorded zero violations across all 359,062 production rows.

The margin construction addresses a production problem. In the calibration ledger, 60% to 95% of reported DSR values equal 0.0. The standard statistic uses a normal CDF, while long-history strategies routinely generate pre-CDF values between minus 40 and minus 15. Retaining that pre-CDF statistic preserves the ranking among failures. Roughly a third of a comparable population received 0.99 or higher under the predecessor in-house score. The current score places 0.002% of rows at or above 0.99.

The evidence has three pieces. A synthetic study covers 2,000 selected strategies with known ground truth, a genuine-edge base rate 0.306, and five years of daily bars at the headline cell. A pre-registered sealed-window test covers 352 single-name strategies from the first 100 alphabetical constituents of the S&P 500 and the S&P MidCap 400, across two model families using 5-minute bars. Fitting ran from 2016-04-19 to 2018-04-19. Evaluation used a sealed 2018-04-20 to 2020-04-20 window under the platform's execution and cost model. The final piece is a ±20% perturbation audit across the full ledger.

We are now building a version of this validation layer and testing it on our own candidates. The copy has limits. Its dispersion constants, correlation matrix and 81 display quantile knots were estimated from 359,062 production records we do not have, so our scale will differ from theirs. We do use the published policy thresholds for the five gates: 0.95, 0.50, 0.10, MinTRL, 0.60. Trial counts, candidate return matrices and cross-validation splits must come from our search as it runs; a market-data table does not contain them, and they cannot be reconstructed afterwards. Our test asks whether the validation layer improves selection among our own candidates. It says nothing about the authors' production population.

Do four extra gates earn their place?

Against the genuine-edge label, the composite records synthetic AUROC 0.9890. Corrected DSR alone reaches 0.9880. The paired-bootstrap gap is +0.0010, with interval [+0.0001, +0.0020], and P(gap greater than zero) is 0.98. Santoni and co-authors treat that result as a tie rather than a lead.

Comparisons with the gates-passed count and the GT-Score proxy are clearer. Their gaps are +0.0292 and +0.0030, and both exceed P of 0.999. Yet corrected DSR alone has lower worst-case regret across the pre-specified 3x2 difficulty grid, 0.009 versus 0.021 for the composite, and it wins the hardest cell outright. The GT-Score proxy, a search objective used here as a familiar benchmark, drops from 0.995 in the easiest cell to 0.806 in the hardest.

Table 1 gives the defence directly: "the composite's case rests on battery coverage and verdict consistency rather than raw discrimination." The coverage exists. Four additional gates catch failure modes invisible to the DSR, while the display guarantee survived all 359,062 rows. A desk ranking its own candidates can still obtain the same ordering from the corrected DSR margin, then inspect the five raw gate values to see the failure.

Four gates for a tenth of a percentage point of AUROC is a reporting decision, and the authors acknowledge it.

The audit reaches the same conclusion from effective contributions. Despite a nominal weight of 0.35, the DSR margin supplies 76.9% of mean absolute contribution across the population. The regime composite contributes 7.6%. At the theoretical perfect point, their order reverses: regime contributes 39.2% on a nominal 0.10, while DSR contributes 32.7% on a nominal 0.35. Nominal weights amount to little more than decoration. The authors make the same observation and cite the composite-indicator literature on the nominal-versus-effective problem.

The 352-strategy null

The abstract reports the null without hiding it. Pooled Spearman between the in-sample raw score and realized out-of-sample Sharpe is 0.013, bootstrap CI [-0.094, 0.121], with one-sided permutation p 0.40. AUROC for positive out-of-sample Sharpe is 0.496. The pre-registered DSR-only baseline returns 0.006. At n=352, the design had roughly 80% power to detect a rank correlation near 0.15, while the confidence interval ends at 0.121.

The authors defend the result in the same sentence by citing limited surviving edge. Median realized out-of-sample Sharpe was -0.15, and only 45% of the 352 strategies finished positive. They also close the most obvious escape route. Cutting the window before 2020-02-19 leaves Spearman at 0.013, median Sharpe at -0.15 and 45% positive. The COVID crash therefore cannot explain the result; the strategies were already losing. An exploratory re-score of 340 strategies with the corrected deflation benchmark produces 0.011, p 0.42 and AUROC 0.503. Old-versus-new predictor rank correlation is 0.983.

The null is honestly reported. As evidence about the score itself, however, it tells us little. The authors put the required remedy first in their future work: repeat the test on a population with quantifiable forward edge, using paper-trading records that let scored strategies accumulate live performance.

Almost unreachable at research horizons

At 300 trials, the history sweep seals genuine-edge searches at rates of 0% after one year, 17% after five years and 73% after twenty. Production is barely looser. Only 3,848 of 359,062 rows pass all five gates, a rate of 1.07%, and every pass belongs to six long-history daily-bar runs. Among 516 run winners, one was sealed.

No strategy sealed in the two-year 5-minute real-market study. The universe-mode cells did find shared parameter sets with Sharpe 1.21 to 2.12 across whole indices, based on 9 to 15 trades in two years. Those results are uncertifiable and deserve to remain so.

Search size moves in the desired direction. For searches containing a genuine candidate, the median raw score declines from 0.37 at 20 trials to 0.016 at 5,000. Over the same range, the winner's median realized Sharpe drops from 1.4 to 0.5. Bigger searches select a genuinely worse winner.

What remains population-specific

The calibration constants belong to the population that produced them, and that population is heavily concentrated. More than 99% of ledger rows come from the top five operators. Run-weighted DSR dispersion is 0.63, compared with the pooled 1.128 adopted by the paper.

PBO dispersion is 10.028. The value suppresses a clamping artifact created by degenerate PBO readings of exactly 0 in 14.4% of rows. The epsilon clamp maps those readings to a logit margin of +13.8. Estimation on the clamp-free core lowers the dispersion to about 1.1, which would move 10% of rows to scores at or above 0.98. The authors describe the choice as deliberate: the larger dispersion keeps the clamped point mass near +1.4 sigma and prevents it from dominating.

The score also lacks a probability interpretation. Expected calibration error is 0.174 at the headline cell, while the offset of 0.5 is an operator default rather than a fitted parameter. Boundary exposure reaches 18.8% of records within 0.05 dispersion units of a threshold. Tightening every threshold by that amount changes the verdict for 311 of 3,848 sealed rows.

Rankings barely move. Worst-case Spearman is 0.9974 under ±20% weight perturbation and exactly 1.0000 when the offset is perturbed.

One result deserves immediate adoption. Empirically estimating deflation variance from sibling search candidates sounds preferable to using a formula, yet it erased discrimination entirely: AUROC fell from 0.96 to 0.50 because sibling Sharpe dispersion mostly reflects estimation noise. We have argued before that a single hold-out window leaves a headline Sharpe exposed to regime luck (our note on Edge Allocation). The search-size sweep is the part we would carry into our own process.

Any reported score should sit beside the five gate values, the search ledger and the Seal. The headline number depends on vintage and population; the five inputs retain their meaning.