AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

A correlation cap tells you nothing about your equal-weight ensemble

Only a persistent IC margin pins the target as PC1, and uniform candidates make that margin exponentially rare.

10 min read

Reviewing: Separated Signal Libraries: Packing, Saturation, and Joint Spectral Limits · Marc da Costa Nunes · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

IC-Margin Separated Multi-Factor Alpha Library

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

Run on a different market than the paper

Not applicable: the paper is formulated for a generic cross-sectional asset universe rather than a specified unavailable market. US equities provide a directly testable instantiation of the cross-sectional signal-library mechanism.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (4)

  • deviation left undescribed by the audit (invalidates: Literal application of the fixed-q uniform product-volume spectrum in equations 10-18 and 42; direct use of the paper's fixed-q numerical admission rates and eigenvalue ratios as empirical predictions)
  • deviation left undescribed by the audit (invalidates: Any attribution of portfolio return, Sharpe ratio, turnover, capacity, or transaction-cost-adjusted performance to the paper)
  • IID evaluation-date assumption for simultaneous IC and pairwise-correlation concentration bounds: The empirical evaluation uses market dates without assuming they are iid. (invalidates: The confidence interpretation and population-cap certification from equation 39, the equation 40 design-eigenvector error guarantee when driven by that bound, and the stated confidence certification from the paper's positive-IC union bound.)
  • Sector-neutral decile portfolio transformation after equal-weight signal aggregation: The traded forecast sector-demeans the continuous EWS, ranks stocks, retains only the highest and lowest deciles, and reweights selected stocks and sector sleeves. (invalidates: Application of Theorem 6.1's eta<rho eigengap and EWS-PC1 angle certificate to the sector-neutral decile portfolio, and any paper-based claim that eta<rho predicts stronger or more stable held-out decile PnL.)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 11.3%Jul 2024
Sharpe
0.11
Total Return
11.3%
Max Drawdown
-39.9%
CAGR
2.4%
Volatility
22.6%
Beta vs SPY
-0.25
Trades
13,259

Automated analysis of the paper

A pairwise correlation ceiling of 0.5 tells you nothing about the direction an equal-weight composite owns. Nunes makes the point in the abstract. He constructs an exponentially large library where every signal has positive information coefficient on every date, all pairs have correlation at most 0.5, and the equal-weight sum is exactly orthogonal to the library's unique leading principal component. The result goes straight at the redundancy screens used before signals enter composites.

A separated signal library admits cross-sectional forecasts sequentially. A new signal enters when its T-date history stays below a correlation cap against every earlier admission. Each accepted signal should also have positive average IC. Equal weighting is expected to retain that average IC as residual structure cancels. The paper tests whether this expectation follows from the admission rule.

This is theory, supported only by synthetic numerical checks. There are no market data. The illustrations use d = 20 assets, hence q = 19 neutral directions, while the tables take T from 1 to 1,000. Its capacity bounds count histories allowed by the geometry. They do not estimate how many predictive signals exist. The exact results instead separate admission rules that force the ensemble mean into the dominant direction from rules that leave ownership unresolved.

There is no Sharpe of theirs anywhere in it.

A library of 3,000 signals remains small

A demeaned, normalized cross-sectional forecast over d assets lies on a sphere with dimension q = d - 1. With 20 assets, that leaves 19 neutral directions, and the ceiling applies to one date. Across T dates, a history occupies a product of T spheres with ambient dimension D = qT. A minimum angle on that product represents the correlation cap on histories.

Using a greedy Gilbert argument on sign words, the paper obtains at least 10^109.9182 mutually 60-degree-separated histories for q = 19 and T = 100. The corresponding count at T = 1,000 is at least 10^1081.8841. For the single-date problem, the relevant object is the kissing number in 19 dimensions. The paper cites Ho's lower bound of 11,948 and a reported numerical semidefinite upper bound of 24,417 from Leijenhorst and de Laat. Interval-arithmetic certification remains unperformed, and the exact value remains undetermined. Once histories replace single cross-sections, the motivating 3,000-signal library is far from the geometric ceiling.

Positive IC also yields a capacity bound, though an exceedingly weak one. Suppose every admitted signal has average IC of at least pbar, the minimum angle is theta, and pbar^2 > cos theta. Then K is at most (1 - cos theta)/(pbar^2 - cos theta). At 60 degrees, corpus-average IC must exceed 1/sqrt(2), about 0.707. With an angle of 85 degrees and pbar = 0.3, the result is K <= 320. It places no useful restriction on the weak-IC regime in which we work.

What does equal weighting own?

The paper's main argument appears in the uniform positive-IC benchmark. Begin with uniform product volume over histories, then retain only histories with positive average IC. Their mean points toward the target. Its target projection is asymptotic to sqrt(2/(pi q T)), about 0.0115 at q = 19 and T = 252 using the paper's own asymptotic. Meanwhile, the uncentered second moment equals I/(qT) exactly. All ambient directions tie, leaving no leading principal component, while the mean falls at T^(-1/2) despite strictly positive average IC for every retained signal.

Centering across designs under this benchmark pushes the intended direction down to the smallest eigendirection. Its eigenvalue is Var(c | c > 0). Near-maximum packing on a fixed domain converges to uniform volume. On the unrestricted sphere, that gives zero mean, an isotropic second moment and a disappearing eigengap. At maximal diversification, equal weighting owns nothing.

Even saturation cannot settle the matter. In the paper's circle example, the upper semicircle contains 3n points spaced pi/(3n) apart and the lower contains 2n, producing 5n points against a maximum of 6n. The imbalance remains 3/5 versus 2/5. As the set covers the circle, its centroid converges to 2/(5 pi), while its second moment stays at I/2 for every n. Coverage therefore coexists with a persistent bias in the equal-weight average.

The margin carries the theorem

A strictly positive IC margin beta changes the spectrum. Under the same uniform product-volume benchmark, now conditioned on average IC of at least beta, the target becomes the unique PC1 at every finite T. No minimum date count is required. As T increases with beta fixed, the EWS norm approaches beta, the leading gap approaches beta^2, and both residual levels fall as 1/T.

The paper gives a coarse sufficient condition, T beta^2 > 1. For beta = 0.03, it first becomes positive at T >= 1,112. Yet uniqueness already holds from T = 1. Treating the sufficient bound as a threshold invents a threshold the result does not have.

The q = 19, beta = 0.03 table needs reading in both directions. From T = 1 to T = 1,000, the target-to-residual eigenvalue ratio rises from 1.117 to 19.022. Across those same rows, admission probability plunges from 0.4500 to 1.77e-5, and the absolute gap declines from 0.00613 to 0.000948. The ratio improves while the quantity estimation error must overcome contracts by a factor of six and a half.

Maintaining a fixed margin costs candidate supply. Under uniform candidates, admission decays as exp(-T I(beta)). At q = 19, beta = 0.1 and T = 100, the refined convolution produces 6.29e-6. The bare exponential term produces 7.49e-5, overstating admission by about an order of magnitude. Candidate pools sized from a large-deviation rate should heed that gap.

Move the margin to the central-limit scale beta = z/sqrt(qT), and admission settles at 15.87% when z = 1. The target-to-residual ratio reaches only 2.525, while the absolute gap vanishes. Nunes replies that rarity under uniform candidates is precisely the point: the rate measures what an actual admission rule must provide, namely a persistent IC margin alongside a distribution of residual structure. Fair enough. The fixed-margin theorems still describe a regime that unguided search will not reach.

Same cap, opposite codebooks

The aligned code gives each signal a common pole with loading rho and balanced sign residuals. Its EWS is the unique history PC1 exactly when rho > 1/(N + 1), where N = (q - 1)T. At a single date, the requirement is rho > 1/q. Take q = 19 and rho = 0.01. The datewise threshold, 1/19, about 0.0526, fails outright. The history threshold holds from T = 6 whenever a balanced code exists at that length. History-level PCA can therefore endorse the composite even as each datewise cross-section rejects it.

The opposing construction uses signals sqrt(0.1) m_t + sqrt(0.3) s v_t + sqrt(0.6) U_t with balanced residual words. Pairwise correlations stay at or below 0.5, and the library grows exponentially at a code rate up to 1 - H_2(5/12) = 0.0201312433 bits per residual coordinate. Set m_t equal to the return target. Every signal then has IC sqrt(0.1) on every date. The mean equals sqrt(0.1) m, while v is the unique PC1, orthogonal to that mean.

Davis-Kahan shows that small moment perturbations preserve the disagreement when operator-norm error remains below half the 0.2 gap. The accompanying remark does not certify continued satisfaction of the pairwise cap after those perturbations. Nunes carefully presents this as an example of what separation allows, rather than a claim about generic positive-IC libraries. He attaches the same limitation to the margin theorems. An arbitrary margin-screened corpus may differ from the uniform conditional law, and Theorem 8.2 retains a fixed positive margin while still disagreeing.

The payoff implication belongs above any signal-review meeting: signal separation has not created PnL separation. In the same construction with m_t = Q_t, every payoff history equals sqrt(0.1) a_t. Variation in return magnitude then makes PnL covariance rank one, with equal leading design weights, even though the leading signal principal component earns zero on every date.

Our US equities run

Because the paper treats a generic cross-section, we implemented our own US-equity version. The paper reports no market results of any kind, leaving none of its figures available for comparison with ours. Our monthly library covers the 1,000 largest non-ADR US names, re-ranked annually by market capitalization. It draws from 28 candidates spanning value, quality, profitability, momentum, reversal, volatility and trading activity.

Each candidate is winsorized at 1/99, demeaned and unit-normalized. The traded book is sector-demeaned and carries a 10% position cap. We consider a candidate when its mean cosine IC across 252 completed next-day pairs reaches at least 0.01. Admission also requires its one-sided mean inner product against every earlier signal to be at most 0.5. Accepted signals are averaged and traded through top and bottom deciles. The portfolio is sector-neutral and dollar-neutral at 100% gross per side, uses MOC fills on the trading day after formation, and holds positions until the next monthly rebalance.

From 2020-01-01 through 2024-07-01, this construction returned 11.30% in total with a Sharpe of 0.11. Annualized volatility was 22.56%, and the drawdown reached 39.86%. Those risk figures dominate the assessment. A sector-neutral, dollar-neutral decile portfolio running 100% gross per side carried far too much risk for the return it delivered.

Our figures include commissions of $0.0040 per share and a $1.00 order minimum. We did not charge short borrow, margin financing or market impact. Zero slippage reflects the auction-print convention used for the MOC fills, rather than an omitted setting.

Two choices in our implementation drive much of the outcome. First, the admission floor of 0.01 is minute on the paper's scale. Using q = 19, T = 252 and beta = 0.01, their qT beta^2 is about 0.48. This lies below the z = 1 case, where the eigenvalue ratio has already fallen to only 2.525 and the absolute gap vanishes. By the paper's arithmetic, our screen had little prospect of creating a dominant target direction. The year-varying universe also departs from the fixed-q product-volume model, so its eigenvalue-ratio benchmarks do not carry over.

Second, we reconstruct the library from empty every month using a fixed insertion order. Candidate order therefore determines admission. The paper proposes randomizing arrivals to measure exactly this order sensitivity. We did not perform that exercise. We also did not report the admitted library's mean energy or residual spectrum, even though the paper calls for those diagnostics. Both omissions are weaknesses in our run, which represents one automated pass rather than a verdict on the authors' work.

What a real library must report

Theorem 6.1 supplies the practical finite-sample criterion without requiring packing optimality or independence. Define rho = ||mu||^2 and let eta be the operator norm of the library's centered second moment across signals. If eta < rho, the leading eigenvalue is simple, with a gap of at least rho - eta. The tangent of the angle between EWS and PC1 is at most eta/(rho - eta).

The identity eta/rho = (1 - rho)/(rho r_eff) shows when additional signals help. They must increase the mean or distribute residual energy more widely. Here r_eff denotes residual effective rank, and it cannot exceed min(K - 1, D).

Estimation determines the true cost of a correlation cap. The paper's uniform bound is eps = sqrt((2/T) log(J(J-1)/alpha)) for a candidate pool fixed before an independent sample of T iid dates. Insert J = 3,000 candidates, T = 252 dates and 95% confidence, and the bound is about 0.39. Certifying a population cap of 0.5 would therefore require an empirical cap near 0.11. Hoeffding is conservative, and the paper describes these as sufficient bounds rather than impossibility results. The direction remains clear: with exponentially many admissible histories and only a year of dates, noise in the selection matrix exceeds the rule's slack. The asymptotic conditions to remember are log J = o(T) and log(J/alpha)/(T theta^4) for a shrinking angle.

Evidence that a real screened library's longitudinal and transverse blocks match the three levels in their Eq. 13 would change my view of the paper's applied relevance. Nunes proposes that comparison as a direct test of the benchmark assumptions. The same construction that meets a fixed margin fails the comparison.

The note has not had independent human peer review and was developed with substantial AI assistance. We have previously shown that a composite may look excellent gross and then lose the comparison after costs in our note on sentiment Black-Litterman. This paper identifies the earlier failure: a composite can appear diversified while owning nothing.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each month-end formation close:
1. Select the year-specific top 1,000 non-ADR US stocks by aiquant_screening_table.capitalization.
2. Compute 28 candidate signals using only prices, daily metrics, and filings available by formation time; apply each stated economic orientation exactly once.
3. On each date-specific eligible panel, winsorize at 1%/99%, demean, and normalize each signal to unit L2 norm.
4. For each candidate, use exactly 252 completed signal/next-trading-day-return target pairs on common panels.
5. Process candidates in fixed insertion order. Retain a candidate for consideration only if mean cosine IC &gt;= 0.01; admit it only if its one-sided mean inner product with every admitted signal is &lt;= 0.5.
6. Average the K admitted normalized signals. If no signal is admitted, do not form a new forecast portfolio.
7. Demean the ensemble forecast within sector and rank the complete eligible cross-section.
8. Select the highest 10% for longs and lowest 10% for shorts. Equal-weight names within each sector and side, then balance sector sleeves so every included sector is net flat; omit sectors lacking both sides.
9. Target 100% long gross and 100% short gross, zero net, subject to a 10% position cap and 4.0 maximum leverage.
10. Execute using actual daily_prices.close on the first trading day after formation. Skip orders lacking that close; never impute an execution or marking price.
11. Hold until the next monthly rebalance, then exit or resize all positions.