Equal-weight admission can favor a forecaster simply because it makes the composite smaller. Soleimani proves the mechanism algebraically and documents it on two real panels where no forecaster aligns with the target. Under his loss, pulling the average toward zero earns a reward even without information about returns.
Alignment, residuals and the admission threshold
The setup will be familiar to anyone who runs cross-sectional books. Each forecaster scores the same units on every date. Those scores are demeaned and scaled to unit norm across units, while the realized return cross-section receives the same treatment. Loss is the squared distance from the composite to that standardized target. For a unit-norm forecast, the loss equals twice one minus cross-sectional correlation. The all-zeros forecast has a score of exactly 1, which the paper names the no-information forecast.
An exact, model-free identity drives the paper. Every standardized forecast has a target alignment, its cross-sectional correlation with the realized target, multiplied by the target. What remains is orthogonal to the target. As a result, the forecast correlation matrix equals the outer product of the alignments plus the covariance matrix of the residuals. For any weight vector, risk separates into an alignment term, the square of (1 minus weighted alignment), and an orthogonal term.
Two implications matter at a trading desk. An equal-weight average of N forecasts can beat the zero forecast only when mean alignment exceeds half of (mean correlation + (1 minus mean correlation)/N). The LLM panel has mean correlation of 0.300 and N of 24, putting the required alignment at about 0.165. Measured mean alignment is 0.006.
The second implication comes from the paper's split between optimally rescaled composite risk and a scale-mismatch penalty. Under weak alignment, that penalty is roughly the squared norm of the composite. Ranking unscaled combinations by risk therefore becomes, to a large extent, a contest to produce the smallest composite.
The empirical analysis uses two panels. Panel A contains 24 language-model forecasters, built from four base models, three personas and two input sets. They rank 60 US large caps each week by forward five-day return across 261 usable dates from January 2021 to June 2026. The test set contains 157 out-of-sample dates. Panel B consists of nine mechanical signals that rank 42 ETFs monthly from 2005 to 2026, including 196 test dates.
Soleimani evaluates candidates with a three-way rule. A candidate enters when a simultaneous heteroskedasticity- and autocorrelation-consistent (HAC) band for its change in pool risk lies entirely below a practical margin. A band above zero means rejection; all other cases remain undecided. The rule runs greedily within nested rolling-origin validation and is compared with peLASSO, exhaustive subset search and ridge-type weights.
Error correlation hides the same dependence
A common screening test asks whether a new model's errors have low correlation with those of the existing book. When every forecast shares one standardized target, deviation correlation is largely a translated version of forecast correlation. At zero alignment, it equals (1 + forecast correlation)/2. This transformation sends [-1, 1] to [0, 1] while cutting dispersion in half.
The LLM panel follows the formula almost perfectly. Its mean forecast correlation is 0.300, while mean deviation correlation is 0.645. The zero-alignment benchmark is 0.650. Across 276 pairs, the relationship is linear with R² of 0.9999. Negative values make up 29% of forecast correlations, yet no deviation correlation is negative. Forecast correlation implies a variance-equivalent ensemble size of 3.03; deviation correlation reduces it to 1.52.
Raw forecast correlation creates another problem. This is the result I would pin above a research desk. Forecasters with better alignment are mechanically more correlated with one another, which makes a search for the "most diverse" candidate favor the least aligned forecast. In the exchangeable-dependence simulations (Design D), mean forecast correlation reverses the candidate ranking. Its Spearman ranges from -0.39 to -0.96 after alignment reaches 0.08. Realized history risk change ranks the same candidates at 0.91 to 0.99.
Under heterogeneous dependence (Design B), forecast correlation has Spearman values from 0.00 to 0.50. The reversal disappears, though the statistic still says little. We discussed a related ordering failure between signal dependence and PnL dependence in an earlier note. Soleimani finds the same trap at the ensemble level.
Six admissions from a shrinking LLM batch
The cross-panel test gives the cleanest example. A batch of value-reversal LLM forecasts has mean correlation of 0.05 with every mechanical rule, leaving it essentially uncorrelated with that pool. The equal-weight procedure admits the batch at all six origins and estimates a risk change of -0.130. The scale-free contribution is -0.0001. Scale mismatch contributes -0.130. Out of sample, the change is -0.114 (t of -9.8). The large t-stat comes from forecasts that reduce the composite's size.
The planted-signal experiment reaches the same conclusion. Zero-loading planted forecasters enter the LLM pool 24% of the time and the signal pool 62% of the time. Inside the LLM panel, the equal-weight three-way rule selects two forecasters at every origin. Their mean within-pool correlation is -0.78.
A strongly anti-correlated pair gets to a small norm quickly.
Soleimani responds with the scale-free rule. It tests only the change in pooled correlation after giving each composite its own history-estimated scale. In Design B, every rule held false admissions to at most 0.1%. The scale-free version sacrifices power for that caution. It leaves 59 to 100% of genuinely improving candidates undecided in the Design B simulations, compared with 10 to 31% under the equal-weight rule.
The dilution bound shows why power vanishes. Suppose 24 incumbents have no alignment. Adding three candidates with alignment 0.10 can increase squared correlation by at most about 0.0005. The paper explicitly describes the scale-free rule as structurally insensitive in large pools. During the positive control, it never admits a candidate into either large pool, even at loading 0.10. Starting selection from scratch works better at loading 0.10: the scale-free rule reaches 77% recovery at precision 0.94, while peLASSO reaches 99% at 0.82.
Every combination remains above the null
The out-of-sample evidence supports the paper's headline: "Selection removes most dilution losses, but no combination beats the no-information forecast." Equal weighting across all 24 LLM forecasts scores 1.292. Three-way selection lowers the score to 1.079, exhaustive search to 1.033 and shrunk ridge weights to 1.019. Even that final result remains +0.019 [0.006, 0.032] above the zero forecast.
The ridge projection finishes at 1.000 because its weights sum to 0.018 and gross exposure is 0.11. In practice, it abstains. Every improvement over equal weighting comes through the orthogonal term, which drops from 0.301 to a range of 0.031 to 0.105. Across the same combinations, the alignment term stays between 0.975 and 0.991.
The ETF panel produces the same pattern. Equal weighting scores 1.283, three-way selection 1.076, exhaustive search 1.061, shrunk affine 1.035 and peLASSO 1.703. Each result exceeds the null, with t from 3.1 to 12.9. The strongest individual signal, 12-1 momentum, records alignment of 0.033 (t of 1.82).
The author states the limits plainly. The equity universe was fixed at the sample's end, and the paper describes the exercise as "a fixed-panel ranking experiment, not an investable backtest." The procedures were also refined on the same data. According to the paper, training-data contamination cannot be ruled out. Two of the four models report cutoffs in March and April 2026, at the end of the sample. Such contamination should make alignment easier to find, leaving a placebo-calibrated p of 0.55 difficult to attribute to it.
Economic diagnostics remain thin. Top-minus-bottom quintile spreads run from -0.10% to 0.50% per five days, and none exceeds the Bonferroni threshold of 3.0. For the three-way pool, the spread is 0.19% per five days with t of 1.0. Transaction costs are absent from the paper. Its sole turnover statistic is origin-to-origin weight turnover of 0.16 for the shrunk ridge weights. The author further writes that the loss "does not constitute a portfolio utility." Bootstrap resampling produces unstable pools, with Jaccard values from 0.35 to 0.49. The risk surface is flat as well: at each origin, 19 to 57 pools fall within 0.005 of the history optimum.
Our separate mechanical run
We could not reproduce the LLM panel. Doing so would require the author's exact model versions, prompts and output collection. We instead built a mechanical cousin of the scale-free rule using 16 technical signals lagged by one session: reversal, momentum, breakout, low volatility, drawdown, gap, intraday reversal and volume surprise.
The signals rank the 300 highest-dollar-volume US stocks and ETFs, screened annually and subject to a $10m daily liquidity gate. Selection uses 504 fit dates and 126 validation dates, then remains fixed for 21-day blocks. Each day at the close, the book buys the top decile and shorts the bottom decile. The run spans 2020-01-01 to 2024-07-01. Every fill incurs commissions of $0.004 a share, with a minimum of $1. Short borrow, financing on up to 4x leverage and market impact are not modelled.
Our test scores a daily next-close decile book in dollars across 300 names. The paper measures relative-score risk on a 60-stock weekly panel and a 42-ETF monthly panel. From 2020-01-01 to 2024-07-01, our book returned 27.20% in total. Its Sharpe was 0.14 (Sortino 0.18, Calmar 0.09), volatility reached 42.65%, and maximum drawdown was -59.82%.
The result is weak. Volatility of 42.65% and a -59.82% drawdown are far outside what a dollar-neutral decile book should produce. The likely sources are the leverage allowance of up to 4x or unequal market sensitivity across the two deciles. A low-volatility signal can create the latter.
The nearest economic figure in the paper is a gross 0.19% five-day spread, at t of 1.0, for the three-way pool. Our figures are a Sharpe of 0.14 and total return of 27.20% over 2020 to 2024, after commissions only. These are different objects: one is a gross five-day quintile spread from a 60-stock weekly panel, while the other is a daily decile book across 300 names. A gap between them provides no evidence about replication.
Our weak result should first be read as evidence from one automated pass. It is not a verdict on the authors' work. The daily horizon, where commissions hurt most, and the conservative scale-free rule are the likeliest explanations. On the LLM panel, that rule retained only its starting forecaster at every tested margin.
When could the zero forecast lose?
An equal-weight combination beats the no-information forecast under this loss only when mean alignment clears the Corollary 1 margin. Across the full panels, equal weighting falls short by 0.159 for LLM and 0.133 for ETF. Even planted forecasters with loading 0.10 produce calibrated selections just 0.027 to 0.031 of risk below the null. If planted alignment begins only in the sample's second half, scale-free recovery drops to 0.25 at loading 0.10.
The decomposition is Soleimani's lasting contribution because it identifies the moving term before anyone pays for the conclusion. My view of the selection machinery would change after seeing a panel whose history alignment clears the margin and whose decile spread, net of borrow, survives the next block.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.