Automated analysis of the paper
A pairwise correlation ceiling of 0.5 tells you nothing about the direction an equal-weight composite owns. Nunes makes the point in the abstract. He constructs an exponentially large library where every signal has positive information coefficient on every date, all pairs have correlation at most 0.5, and the equal-weight sum is exactly orthogonal to the library's unique leading principal component. The result goes straight at the redundancy screens used before signals enter composites.
A separated signal library admits cross-sectional forecasts sequentially. A new signal enters when its T-date history stays below a correlation cap against every earlier admission. Each accepted signal should also have positive average IC. Equal weighting is expected to retain that average IC as residual structure cancels. The paper tests whether this expectation follows from the admission rule.
This is theory, supported only by synthetic numerical checks. There are no market data. The illustrations use d = 20 assets, hence q = 19 neutral directions, while the tables take T from 1 to 1,000. Its capacity bounds count histories allowed by the geometry. They do not estimate how many predictive signals exist. The exact results instead separate admission rules that force the ensemble mean into the dominant direction from rules that leave ownership unresolved.
There is no Sharpe of theirs anywhere in it.
A library of 3,000 signals remains small
A demeaned, normalized cross-sectional forecast over d assets lies on a sphere with dimension q = d - 1. With 20 assets, that leaves 19 neutral directions, and the ceiling applies to one date. Across T dates, a history occupies a product of T spheres with ambient dimension D = qT. A minimum angle on that product represents the correlation cap on histories.
Using a greedy Gilbert argument on sign words, the paper obtains at least 10^109.9182 mutually 60-degree-separated histories for q = 19 and T = 100. The corresponding count at T = 1,000 is at least 10^1081.8841. For the single-date problem, the relevant object is the kissing number in 19 dimensions. The paper cites Ho's lower bound of 11,948 and a reported numerical semidefinite upper bound of 24,417 from Leijenhorst and de Laat. Interval-arithmetic certification remains unperformed, and the exact value remains undetermined. Once histories replace single cross-sections, the motivating 3,000-signal library is far from the geometric ceiling.
Positive IC also yields a capacity bound, though an exceedingly weak one. Suppose every admitted signal has average IC of at least pbar, the minimum angle is theta, and pbar^2 > cos theta. Then K is at most (1 - cos theta)/(pbar^2 - cos theta). At 60 degrees, corpus-average IC must exceed 1/sqrt(2), about 0.707. With an angle of 85 degrees and pbar = 0.3, the result is K <= 320. It places no useful restriction on the weak-IC regime in which we work.
What does equal weighting own?
The paper's main argument appears in the uniform positive-IC benchmark. Begin with uniform product volume over histories, then retain only histories with positive average IC. Their mean points toward the target. Its target projection is asymptotic to sqrt(2/(pi q T)), about 0.0115 at q = 19 and T = 252 using the paper's own asymptotic. Meanwhile, the uncentered second moment equals I/(qT) exactly. All ambient directions tie, leaving no leading principal component, while the mean falls at T^(-1/2) despite strictly positive average IC for every retained signal.
Centering across designs under this benchmark pushes the intended direction down to the smallest eigendirection. Its eigenvalue is Var(c | c > 0). Near-maximum packing on a fixed domain converges to uniform volume. On the unrestricted sphere, that gives zero mean, an isotropic second moment and a disappearing eigengap. At maximal diversification, equal weighting owns nothing.
Even saturation cannot settle the matter. In the paper's circle example, the upper semicircle contains 3n points spaced pi/(3n) apart and the lower contains 2n, producing 5n points against a maximum of 6n. The imbalance remains 3/5 versus 2/5. As the set covers the circle, its centroid converges to 2/(5 pi), while its second moment stays at I/2 for every n. Coverage therefore coexists with a persistent bias in the equal-weight average.
The margin carries the theorem
A strictly positive IC margin beta changes the spectrum. Under the same uniform product-volume benchmark, now conditioned on average IC of at least beta, the target becomes the unique PC1 at every finite T. No minimum date count is required. As T increases with beta fixed, the EWS norm approaches beta, the leading gap approaches beta^2, and both residual levels fall as 1/T.
The paper gives a coarse sufficient condition, T beta^2 > 1. For beta = 0.03, it first becomes positive at T >= 1,112. Yet uniqueness already holds from T = 1. Treating the sufficient bound as a threshold invents a threshold the result does not have.
The q = 19, beta = 0.03 table needs reading in both directions. From T = 1 to T = 1,000, the target-to-residual eigenvalue ratio rises from 1.117 to 19.022. Across those same rows, admission probability plunges from 0.4500 to 1.77e-5, and the absolute gap declines from 0.00613 to 0.000948. The ratio improves while the quantity estimation error must overcome contracts by a factor of six and a half.
Maintaining a fixed margin costs candidate supply. Under uniform candidates, admission decays as exp(-T I(beta)). At q = 19, beta = 0.1 and T = 100, the refined convolution produces 6.29e-6. The bare exponential term produces 7.49e-5, overstating admission by about an order of magnitude. Candidate pools sized from a large-deviation rate should heed that gap.
Move the margin to the central-limit scale beta = z/sqrt(qT), and admission settles at 15.87% when z = 1. The target-to-residual ratio reaches only 2.525, while the absolute gap vanishes. Nunes replies that rarity under uniform candidates is precisely the point: the rate measures what an actual admission rule must provide, namely a persistent IC margin alongside a distribution of residual structure. Fair enough. The fixed-margin theorems still describe a regime that unguided search will not reach.
Same cap, opposite codebooks
The aligned code gives each signal a common pole with loading rho and balanced sign residuals. Its EWS is the unique history PC1 exactly when rho > 1/(N + 1), where N = (q - 1)T. At a single date, the requirement is rho > 1/q. Take q = 19 and rho = 0.01. The datewise threshold, 1/19, about 0.0526, fails outright. The history threshold holds from T = 6 whenever a balanced code exists at that length. History-level PCA can therefore endorse the composite even as each datewise cross-section rejects it.
The opposing construction uses signals sqrt(0.1) m_t + sqrt(0.3) s v_t + sqrt(0.6) U_t with balanced residual words. Pairwise correlations stay at or below 0.5, and the library grows exponentially at a code rate up to 1 - H_2(5/12) = 0.0201312433 bits per residual coordinate. Set m_t equal to the return target. Every signal then has IC sqrt(0.1) on every date. The mean equals sqrt(0.1) m, while v is the unique PC1, orthogonal to that mean.
Davis-Kahan shows that small moment perturbations preserve the disagreement when operator-norm error remains below half the 0.2 gap. The accompanying remark does not certify continued satisfaction of the pairwise cap after those perturbations. Nunes carefully presents this as an example of what separation allows, rather than a claim about generic positive-IC libraries. He attaches the same limitation to the margin theorems. An arbitrary margin-screened corpus may differ from the uniform conditional law, and Theorem 8.2 retains a fixed positive margin while still disagreeing.
The payoff implication belongs above any signal-review meeting: signal separation has not created PnL separation. In the same construction with m_t = Q_t, every payoff history equals sqrt(0.1) a_t. Variation in return magnitude then makes PnL covariance rank one, with equal leading design weights, even though the leading signal principal component earns zero on every date.
Our US equities run
Because the paper treats a generic cross-section, we implemented our own US-equity version. The paper reports no market results of any kind, leaving none of its figures available for comparison with ours. Our monthly library covers the 1,000 largest non-ADR US names, re-ranked annually by market capitalization. It draws from 28 candidates spanning value, quality, profitability, momentum, reversal, volatility and trading activity.
Each candidate is winsorized at 1/99, demeaned and unit-normalized. The traded book is sector-demeaned and carries a 10% position cap. We consider a candidate when its mean cosine IC across 252 completed next-day pairs reaches at least 0.01. Admission also requires its one-sided mean inner product against every earlier signal to be at most 0.5. Accepted signals are averaged and traded through top and bottom deciles. The portfolio is sector-neutral and dollar-neutral at 100% gross per side, uses MOC fills on the trading day after formation, and holds positions until the next monthly rebalance.
From 2020-01-01 through 2024-07-01, this construction returned 11.30% in total with a Sharpe of 0.11. Annualized volatility was 22.56%, and the drawdown reached 39.86%. Those risk figures dominate the assessment. A sector-neutral, dollar-neutral decile portfolio running 100% gross per side carried far too much risk for the return it delivered.
Our figures include commissions of $0.0040 per share and a $1.00 order minimum. We did not charge short borrow, margin financing or market impact. Zero slippage reflects the auction-print convention used for the MOC fills, rather than an omitted setting.
Two choices in our implementation drive much of the outcome. First, the admission floor of 0.01 is minute on the paper's scale. Using q = 19, T = 252 and beta = 0.01, their qT beta^2 is about 0.48. This lies below the z = 1 case, where the eigenvalue ratio has already fallen to only 2.525 and the absolute gap vanishes. By the paper's arithmetic, our screen had little prospect of creating a dominant target direction. The year-varying universe also departs from the fixed-q product-volume model, so its eigenvalue-ratio benchmarks do not carry over.
Second, we reconstruct the library from empty every month using a fixed insertion order. Candidate order therefore determines admission. The paper proposes randomizing arrivals to measure exactly this order sensitivity. We did not perform that exercise. We also did not report the admitted library's mean energy or residual spectrum, even though the paper calls for those diagnostics. Both omissions are weaknesses in our run, which represents one automated pass rather than a verdict on the authors' work.
What a real library must report
Theorem 6.1 supplies the practical finite-sample criterion without requiring packing optimality or independence. Define rho = ||mu||^2 and let eta be the operator norm of the library's centered second moment across signals. If eta < rho, the leading eigenvalue is simple, with a gap of at least rho - eta. The tangent of the angle between EWS and PC1 is at most eta/(rho - eta).
The identity eta/rho = (1 - rho)/(rho r_eff) shows when additional signals help. They must increase the mean or distribute residual energy more widely. Here r_eff denotes residual effective rank, and it cannot exceed min(K - 1, D).
Estimation determines the true cost of a correlation cap. The paper's uniform bound is eps = sqrt((2/T) log(J(J-1)/alpha)) for a candidate pool fixed before an independent sample of T iid dates. Insert J = 3,000 candidates, T = 252 dates and 95% confidence, and the bound is about 0.39. Certifying a population cap of 0.5 would therefore require an empirical cap near 0.11. Hoeffding is conservative, and the paper describes these as sufficient bounds rather than impossibility results. The direction remains clear: with exponentially many admissible histories and only a year of dates, noise in the selection matrix exceeds the rule's slack. The asymptotic conditions to remember are log J = o(T) and log(J/alpha)/(T theta^4) for a shrinking angle.
Evidence that a real screened library's longitudinal and transverse blocks match the three levels in their Eq. 13 would change my view of the paper's applied relevance. Nunes proposes that comparison as a direct test of the benchmark assumptions. The same construction that meets a fixed margin fails the comparison.
The note has not had independent human peer review and was developed with substantial AI assistance. We have previously shown that a composite may look excellent gross and then lose the comparison after costs in our note on sentiment Black-Litterman. This paper identifies the earlier failure: a composite can appear diversified while owning nothing.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.