Signal correlation cannot tell a trading desk which pair of alphas will have the lower PnL correlation. Nunes proves the stronger statement, in his words: "neither average signal similarity nor PnL correlation determines, nontrivially bounds, or universally orders the other." The proof is elementary. Its implications for screening candidate alphas are harder to dismiss.
The geometry behind the result
Everything starts with two correlations taken over different index sets. Let there be $d$ assets on dates $t$, forward returns $R_t$, and cross-sectional signals $S_{i,t}$ demeaned and scaled to unit Euclidean norm. Separate each return into a market component and a demeaned component. Write the magnitude of the demeaned piece as $a_t$ and its direction as $Q_t$. The signal's inner product with $Q_t$ is its realized cross-sectional Pearson IC. Trading the signal vector produces PnL equal to $a_t$ times that IC. Qian and Hua (2004) give the same single-period identity. Nunes casts it in realized, Euclidean-normalized form.
The decomposition follows directly. For any date, the correlation between two signals is the product of their ICs plus a residual inner product. The residual contains whatever remains after the return direction has been projected out. Divide it by its attainable radius and it becomes exactly the cross-sectional partial correlation between the signals, controlling for the realized return vector.
Time averaging separates signal correlation into an uncentered IC cross-moment and a transverse term. The transverse term contains everything orthogonal to the return direction. Pearson PnL correlation discards that geometry and uses only the centered IC series, weighted by $a_t^2$.
The amount discarded is enormous. With a realized IC of 0.03, the transverse norm is 0.99955.
99.91% of the squared signal norm generated no PnL that day.
A library screen based on signal correlation is therefore dominated by geometry orthogonal to the day's realized return direction. The paper's Theorem 9 turns that observation into a short construction. When $d \geq 4$, the transverse space has at least two dimensions. Assign two signals the same longitudinal coordinate on every date, giving them identical exposure along $Q_t$. Their PnLs match exactly, and PnL correlation is 1. Yet their date-$t$ signal similarity is $p_t^2 + (1-p_t^2)\langle V_t, W_t\rangle$. Because the last inner product may take any value in $[-1,1]$, signal correlation remains unrestricted.
The reverse construction works too. Choose a target signal similarity $c$, then take two longitudinal series bounded in magnitude by $\varepsilon$, with time-series correlations of 0 and 0.9. A transverse completion exists whenever $|c| \leq 1 - 2\varepsilon^2$. With $d = 3$, the transverse space is one-dimensional, leaving only a weaker discrete result. The paper describes this as an exact sharpening of the simulation result in Sorensen, Qian, Schoen and Hua (2004).
Markets do not enter the evidence. The paper makes no empirical claim about them and contains no real data of any kind. The author writes that the synthetic experiment "should eventually be replaced or supplemented by an out-of-sample study on a real signal library."
When the orthogonal component matters
For a single signal traded at fixed Euclidean norm, the transverse component costs nothing. By construction, it contributes nothing to that date's payoff. Once signals are combined and the result is renormalized, however, the transverse Gram matrix returns through the denominator.
The signal Gram matrix is $A_t = p_t p_t' + U_t'U_t$. Holding the individual ICs fixed, greater transverse similarity increases the combination's norm and reduces its normalized IC. Transverse cancellation raises it. Weights proportional to $A_t^{-1}p_t$ produce the best Euclidean-normalized combination. The value attained is the multiple correlation between the realized return direction and the signal span.
Under a no-information null, the squared multiple correlation follows a Beta distribution with mean $K/(d-1)$. This null takes a rank-$K$ span independent of an isotropic $Q_t$. At $K = 12$ and $d = 500$, its mean is 0.0240. The simulation produces a realized mean of 0.0334 and an adjusted value of 0.0096. When the IC innovation standard deviation is 0.012, the mean drops to 0.0055, below the null, while the adjusted value reaches $-0.0190$.
Raw ex post span $R^2$ can rise mechanically as the library grows. It may also fall below the null when the assigned per-date ICs are smaller than random orientation would imply.
Both optimization rules are ex post envelopes because $p_t$ includes the realized forward return. This applies to the Euclidean solution $A_t^{-1}p_t$ and the risk-metric solution $B_t^{-1}p_t$, where $B_t = S_t'\Sigma_t S_t$. Look-ahead is built into both span-efficiency diagnostics, and Nunes states that plainly. A useful side result follows: the rules coincide for every longitudinal exposure if and only if the risk metric is proportional to the Euclidean form on the signal span.
An implementer gets three figures for each pair: average signal similarity $C_{sig}$, the IC cross-moment $C_{IC}$, and the difference between them. The transverse statistic also comes with a reference scale, $t$ with $d-3$ degrees of freedom and Fisher-z variance $1/(d-4)$. This law requires normalized transverse directions that are independent and uniform on the sphere. The paper acknowledges the weakness directly: "the isotropic/Gaussian null is not a realistic generative model for signals that come out of a shared feature pipeline." It supplies no rule for converting the three figures into weights before returns are observed.
A simulation designed to separate the two
The simulation contains $d = 500$, $T = 1000$, $K = 12$ signals and 66 pairs, using seed 20260906. Design alphas range from 0.015 to 0.025. The IC innovation standard deviation is 0.045, deliberately set near the $1/\sqrt{d-1} = 0.0448$ single-date IC noise floor.
Across the 66 pairs, average signal similarity has a Spearman rank correlation of $-0.070$ with PnL correlation. The corresponding figure for the IC cross-moment is 0.974. At innovation scales 0.012, 0.025 and 0.070, the first relationship measures $-0.165$, $-0.160$ and $+0.150$; the second measures 0.388, 0.877 and 0.989. Some pairs make the separation vivid. Pair 3-12 records signal correlation of 0.4766 alongside PnL correlation of $-0.405$, with almost the entire signal relationship sitting in the transverse component.
The design creates this gap deliberately. Signals belong to three feature families, and their loadings cycle through 0.15, 0.35, 0.60 and 0.80. Median absolute transverse similarity is 0.1641 within a family and 0.00094 across families. IC dependence comes from an entirely separate factor. Nunes says exactly what this means: "The independence is imposed, not discovered; whether it holds approximately in a real signal library is an empirical question."
One seed. One family and dispersion configuration. Only the IC innovation scale changes, taking values of 0.012, 0.025, 0.045 and 0.070.
There is no repeated-seed variability and there are no confidence intervals. Section 11 recommends block bootstraps and false-discovery-rate control for the pairwise table, but neither procedure is applied to the 66 pairs. Appendix A prints the complete NumPy generator, making the design reproducible. The acknowledgments disclose that an AI system executed the experiment as a check, with no independent human replication reported.
Those limitations leave the theorem untouched because it requires no data. They do constrain the ranking claim. The evidence that IC co-alignment is a better diversification screen than signal correlation comes from a construction.
Our allocator
We used the paper's normalization and its IC-plus-PnL covariance logic to build a dollar-neutral US equity book. Daily bars cover 2020-01-01 through 2024-07-01. The universe is the annual point-in-time top 500 by capitalization, excluding ADRs and requiring positive lagged 20-day median dollar volume.
The book has four signals: 12-month momentum skipping 21 days, short-term reversal over five days, value based on at least three of four components, and quality based on at least four of six. Each signal is winsorized at 1%/99% and residualized on an intercept plus log market cap. We then demean it and scale it to unit Euclidean norm, following the paper's convention.
Weights use a trailing 252-day window with a minimum of 126 days. The objective trades expected signed payoff against a half-and-half blend of IC-scale variance and implemented-PnL variance. We apply diagonal shrinkage of 25% and risk aversion of 5.0. An L1 turnover penalty of 0.10 and a 0.50 limit on each signal weight complete the optimization. Security weights are demeaned, normalized to 1.0 gross, limited to 10% absolute and executed MOC at $0.0040 per share.
Three modeling decisions belong to us. The paper's optimum uses realized $p_t$; we replaced it with a trailing signed IC estimate, so our weights are not the paper's optimum. Its risk term is the single-date cross-sectional Gram $S_t'\Sigma_t S_t$, whereas ours blends that quantity with time-series payoff variance. The 10% cap and gross normalization also break the identity $\Pi = a_t \times \mathrm{IC}$ outright. The paper identifies that identity as a scope condition. In our implementation, the weight norm becomes a third multiplicative channel.
From 2020-01-01 to 2024-07-01, the resulting book returned $-17.65\%$ in total. Its Sharpe was $-0.83$, maximum drawdown was $-18.12\%$, and annualized volatility was 4.72%. These figures describe our build over four and a half years. The paper gives no real-data performance result, leaving nothing for us to reproduce or beat.
With four signals and $d = 500$, the isotropic-null span mean is $4/499 = 0.0080$. This library is tiny beside the diagnostic's own noise floor. We charged commissions on every fill. Short borrow on the short leg, financing and market impact are not included.
Evidence that would make it tradable
The decomposition is correct, and the theorem is airtight. I would still hold off on the ranking claim. A real library of forty or fifty signals would change that view. Compute all $K(K-1)/2$ pairs and attach block-bootstrap intervals. If same-family absolute transverse similarity separates from cross-family similarity as clearly as 0.1641 against 0.00094, the diagnostic deserves a permanent place in our screening.
We have previously argued that the verdict from a correlation figure can depend on the choices around it rather than the data (our ESGU note). This paper identifies the relevant choice and provides an isotropic reference scale against which to test it.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.