AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

The isotropic threshold binds at 20 assets and vanishes at 500

Nunes gives an exact alignment test and no market data. We built the ensemble on 500 US stocks.

2026-09-16 · 8 min read · US equities

Reviewing: Large Signal Libraries: Equal-Weight Limits and the Divergent Spectra of Signals and PnL · Marc da Costa Nunes · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Large-Library Dollar-Neutral Cross-Sectional Multi-Factor Ensemble

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

Run on a different market than the paper

The paper's formal results are asset-agnostic and concern the geometry of normalized cross-sectional signal portfolios rather than a market-specific instrument or futures-curve mechanism. The implementation would apply the same signal-library aggregation and PCA diagnostics to a liquid US-equity cross-section; reported motivating observations, if any, do not transfer to this universe.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (9)

  • Maximum empirical library size and nested growth path: The repaired specification contains at most 1461 distinct signals and drops the paper's proposed K=2000 and K=3000 empirical-program settings. (invalidates: The paper's proposed empirical growth-path diagnostics at K=2000 and K=3000 and any application of the motivating 3000-signal K/q ratio to this implemented library.)
  • Random seed for empirical library orderings: The empirical ordering procedure reuses seed 20260911. (invalidates: Any claim that the empirical ordering paths or their randomization distribution are reproduced from a paper-prescribed random seed.)
  • Ledoit-Wolf spectral regularization: The specification mandates Ledoit-Wolf linear shrinkage for finite-history covariance and correlation spectra. (invalidates: Any attribution of the resulting finite-history covariance or correlation eigenvalues, eigengaps, eigenvectors, and ensemble-PC alignment estimates to a uniquely paper-specified estimator.)
  • Family-balanced growth path above K=200: Exact without-replacement family balance is evaluated only at K=25,50,100,200; it is unavailable at K=500,1000,1461. (invalidates: The paper's proposed balanced-over-families growth-path comparisons at K=500, K=1000, and the implemented maximum K=1461.)

5 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total -29.7%Jul 2024
Sharpe
-0.16
Total Return
-29.7%
Max Drawdown
-97.0%
CAGR
-7.6%
Volatility
41.1%
Beta vs SPY
-0.54
Trades
554,328

Does 158 signals per dimension explain anything?

Three thousand cross-sectional signals over twenty assets occupy nineteen dimensions after demeaning, roughly 158 submissions for every available direction. Nunes considers whether this crowding makes an equal-weight ensemble from a large library resemble the leading return factor. He rejects the explanation. In his formulation, $K/q$ captures algebraic crowding. It says nothing about the shape or mean of the sampling measure.

Consider iid uniform points on the sphere $S^{q-1}$. Their population mean is zero, while their second moment remains $I/q$ for every $K$. Averaging any number of them therefore approaches zero, with no unique leading direction available for alignment. The opposite degenerate case has every submission copying one design. From $K=1$, the ensemble equals that design and its second moment is rank one from $K=1$. Mean-PC1 agreement is perfect, although the library has neither diversification nor a new direction.

Exact alignment between an equal-weight library and its own leading principal component can signal duplication just as readily as structure. Nunes begins there, correctly. The remaining question is why anyone treated the motivating 90% as a consequence of library size.

Three thousand designs, one book

The large signal library represents a research process that repeatedly submits cross-sectional portfolios over the same assets. Demeaning each submission makes it dollar neutral, and normalization gives it unit length. Equal weights then combine the library into one book. The paper studies the limit of that book as submissions accumulate.

Demeaning removes one dimension. For $d$ assets, signals occupy a $q = d-1$ dimensional neutral space and, after normalization, lie on the sphere $S^{q-1}$. Twenty assets means $q=19$ and $S^{18}$. The 3,000 by 3,000 signal Gram matrix at every date consequently has rank at most 19.

The PnL relation is a familiar identity. Express the demeaned realized return as its magnitude multiplied by a unit direction. The projection of a normalized signal onto that direction is its realized cross-sectional IC. PnL for the date equals dispersion times IC. Orthogonal components contribute nothing.

Nunes supplies population theory and one synthetic sweep. With iid draws from a fixed design distribution, the equal-weight ensemble converges to the distribution's mean. Exchangeability changes the limit to a random conditional mean $E[S_{Z_1}|\mathcal{G}]$. This describes a research group whose shared environment produces an ensemble that is stable yet environment-dependent. The paper also gives dependence a price. For equicorrelated designs, averaging error is $\sigma^2(\rho + (1-\rho)/K)$, and the effective count $K/(1+(K-1)\rho)$ converges to $1/\rho$. Pairwise correlation of 0.2 leaves an effective count of five signals, regardless of the number submitted.

The motivating observation arrives second-hand. An equal-weight ensemble of about 3,000 signals across 20 assets reportedly had roughly 90% correlation with the leading component of the asset-space return structure, identified in that market as reversal. The account gives no period, universe identity, frequency, or estimation window. The paper states the limitation in its abstract: no market-data empirical results are presented.

Its main service is keeping four objects separate. The equal-weight signal is a first moment. Signal principal components are another object. Equal-weight PnL and PnL principal components follow, with the latter calculated after projection onto returns, dispersion weighting and time-series centering. Signal PCA divides again between a date-level operator on the asset space and a whole-history operator on signal processes. Return PCA is a fifth object because it averages dates rather than designs. Before a reported 90% correlation can be interpreted, the reader needs to know which object was computed, whether time-series centering was applied, and the window used.

5.26% at twenty assets, 0.20% at five hundred

The clean theorem concerns a design population invariant under orthogonal transformations that fix the mean direction. Under that assumption, the mean is the unique leading eigenvector of the uncentered signal second moment exactly when longitudinal energy $E[(m^\top S)^2]$ exceeds the isotropic share $1/q$. The equivalent condition requires longitudinal variance plus mean pairwise signal correlation to cross the same threshold.

For 20 assets, $q=19$, putting the threshold at 5.26%. Five and a quarter percent is a substantial hurdle for a signal library, which makes the motivating case worth examining. With $d=500$, $q=499$ and the threshold shrinks to 0.2004%. Almost any cloud with faint longitudinal concentration clears it. Provided that concentration does not decline at the same rate as $1/q$, the same alignment statement moves from informative to nearly vacuous as universe size rises. At a 0.2004% threshold, subject to that same concentration caveat, signal-PC1 alignment in a broad US cross-section says very little.

The sole numerical example is synthetic: axial clouds on $S^{18}$, $K = 3000$ unit vectors, 146 clouds, seed 20260911, with $19a^2$ swept from about 0.5 to 3.0. The mean becomes PC1 exactly when $qa^2 > 1$. At $q=19$, each of the 18 transverse eigenvalues is $(1-a^2)/18$. Population geometry changes sharply at the threshold, while finite-sample alignment shifts across a neighborhood around it. Nunes describes the run as "one reproducible sample path, not a confidence band or a return-data experiment."

The assumption carrying the result

Axial symmetry describes how the research process filled the sphere. Nunes offers one diagnostic: axial symmetry implies a flat transverse spectrum. Mean-PC1 alignment by itself cannot establish the assumption.

The result without symmetry is more useful. First compute the transverse residual of the mean direction. Then compare the direction's own energy with the largest transverse eigenvalue. Their ratio bounds the sine of the angle to sample PC1. This Rayleigh-quotient residual estimate applies directly to the observed matrix. A non-positive gap leaves it silent, and a small residual near degeneracy certifies nothing.

Two further bounds are quantitative. Combining design weights into signals contracts the tangent of the angle to PC1 by at most $\sqrt{\lambda_2/\lambda_1}$, with a sharp factor: equality holds for $u = \alpha v_1 + \beta v_2$ with $\alpha\beta \neq 0$ and $\alpha^2 + \beta^2 = 1$. The empirical leading eigendirection also lies within $2|A_{\mathrm{sig},K} - A_{\mathrm{sig}}|_{\mathrm{op}}/\delta$ in sine of the population direction, where $\delta = \lambda_1 - \lambda_2 > 0$. In practice, both bounds depend on the gap one can estimate.

PnL sees one number per date

Two libraries may share identical IC processes while having arbitrarily different transverse components. Their PnL kernels remain identical, even though their signal principal components can differ without bound. A signal-correlation matrix based on cross-sectional overlap may therefore contain structure that never produces PnL. Risk budgeting that substitutes one object for the other carries the same mistake. We took this position when reviewing Nunes's earlier decomposition of signal correlation and PnL dependence (our note). The present paper develops that argument at the operator level and follows it as the library grows.

Principal components can be reordered by discarding the transverse component or by transforming the longitudinal component, even when every population moment is known exactly. The longitudinal transformation combines two operations: quadratic reweighting of dates by return dispersion and temporal centering. Uncentered and centered signal PCA have the same eigenvector compatibility condition, yet the leading eigenvalue differs by the squared mean norm. A direction that leads uncentered PCA may lose that position after centering.

Who checked the 90%?

The paper's proposed empirical program is its defence. It calls for nested library sizes $K = 25, 50, 100, 200, 500, 1000, 2000, 3000$, with randomized orderings balanced across research families. Row-degree dispersion would check equal centrality. Eigengaps would appear with resampling uncertainty. A Marchenko-Pastur reference edge $\sigma^2(1+\sqrt{\gamma})^2$ would be used only where its assumptions apply. The program is reasonable within those limits.

When library size and history length increase together, the paper treats the joint regime as a high-dimensional spectral problem that requires regularization and noise calibration. It concedes that the temporal estimation term "need not vanish in operator norm." Empirical PnL covariance has rank no greater than $\min(K, T-1)$. With $T$ dates of history, that permits at most $T-1$ nonzero directions for $K = 3{,}000$ designs. The asymptotics also leave untouched the selection bias of a library assembled through thousands of submissions.

The 90% was never re-measured.

We ran it on 500 US stocks and lost 29.74%

The results in the paper are asset-agnostic, and its motivating universe of 20 unnamed assets could not be traded by us. We substituted a liquid US equity cross-section for that setting. Because this is a theory paper with no market-data results, our test is a construction based on its idea and does not replicate a claim made by the authors.

Our library grew to 1,461 distinct designs across five families using the annual top 500 US stocks by capitalization. It contained 432 momentum, 600 reversal, 45 volatility, 240 volume and 144 point-in-time fundamentals designs. Names had to pass a one-day-lagged 63-day median dollar volume screen at $1,000,000. Every cross-section was winsorized at 1%/99%, demeaned, normalized and equally averaged. We capped the neutral vector at 10% per name, demeaned it again after clipping, and scaled it to gross leverage no greater than 4.0. Signals through the $t$ close traded daily at the $t+1$ close using observed closes. Costs were $0.0040$ per share with a $1.00$ minimum.

The backtest window ran from 2020-01-01 to 2024-07-01. Our run lost 29.74% in total during that period. Sharpe was -0.16, Sortino -0.20 and Calmar -0.08. Maximum drawdown reached -96.99%, with annualized volatility of 41.06%.

That 41.06% annualized volatility reflects the construction as well as the market. Dollar neutrality and a 4.0 gross cap determine scale, so a lower cap would shift the entire strip. Our cost model excludes short borrow and margin financing on up to four times capital. In a daily-rebalanced neutral book of this size, the omission is the likeliest gap between the -96.99% drawdown and a tradeable result. This was one automated pass built from the paper's description, rather than a verdict on the authors' work. The weak outcome applies to our implementation choices, especially the $d=500$ universe. There, the paper's own threshold drops to 0.2004%, leaving its geometry with almost no constraint on the result.

The mathematics holds up. The opening 90% remains untested. Running the 20-asset case with a residual-and-gap certificate at every date, then reporting the gap and its resampling spread, would make the motivating claim checkable. Until someone does that, 158 signals per dimension measures crowding and nothing else.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

1. On each date, select the annual top-500 STOCK universe by
   aiquant_screening_table.capitalization; exclude ADRs.
2. Require a one-day-lagged 63-day rolling median dollar volume of at least
   $1,000,000. Exclude assets with missing liquidity data.
3. Build up to 1,461 distinct designs:
   - 432 price-momentum
   - 600 price-reversal
   - 45 volatility
   - 240 volume
   - 144 point-in-time fundamental
4. For fundamentals, admit records only when datekey <= signal as-of date;
   never use calendardate as the availability field.
5. On one canonical common asset panel for the complete realized dictionary,
   winsorize each cross-section at 1%/99%, demean it, and normalize it to unit
   Euclidean norm. Do not rebuild the panel or renormalize when K changes.
6. Form the equal-weight ensemble m_K(t) = (1/K) * sum_i S_i(t), using the
   largest valid nested library up to K=1,461. Keep this average distinct from
   signal PC1, PnL PC1, and return PC1.
7. Convert m_K into a dollar-neutral long-short portfolio. Cap each position at
   10% of capital, re-demean after clipping, and uniformly scale gross leverage
   to no more than 4.0. Skip orders if a valid neutral vector cannot be formed.
8. Signals observed through date t close are implemented at date t+1 close
   using an observed daily_prices.close. If that execution price is missing,
   skip the affected trade; never synthesize or forward-fill a fill price.
9. Rebalance on the configured daily schedule; the specification also defines
   a weekly Friday/last-trading-day schedule for separate evaluation.
10. Monthly, estimate signal and payoff geometry over the preceding 504 valid
    trading days. Report uncentered and centered signal spectra, longitudinal
    and transverse kernels, Ledoit-Wolf-shrunk PnL covariance/correlation
    spectra, effective-count diagnostics, concentration, rank, and eigengaps.