Equation (3) forces anyone implementing this paper to make a choice the text leaves unresolved.
First, the construction.
The portfolio Molero Gonz\filez and coauthors built
The authors take daily log returns for 484 S&P 500 names across 5,029 trading days, from 6 January 2005 to 21 November 2024. They estimate the correlation matrix over a rolling 504-day window, advancing it 60 days each time. The procedure produces 76 windows.
After normalisation, each eigenvector becomes a portfolio. The largest is the market mode. As a proxy, the authors use an equally weighted 1/p portfolio of the 484 names, with mean daily log return 0.0390%, standard deviation 1.2919%, skewness -0.6056 and a worst day of -13.58%.
The paper focuses on the second eigenvector. Earlier work by the same group found that the second-largest eigenvalue moves against the market mode under stress, making its portfolio a possible defensive overlay. The authors first remove the market through an OLS single-index regression of each stock on the market factor within each window. They then calculate the correlation matrix from the residuals. Its largest eigenvector, v1, serves as a stand-in for the second raw eigenvector.
The match is close. Pearson correlation stays above 0.95 through most of the sample. It falls between 0.90 and 0.95 from August 2010 to September 2012, and drops below 0.85 only during January to September 2007 and September 2022 to June 2023.
N_eff is the paper's new contribution. It is built from the Inverse Participation Ratio, used in finance since at least Utsugi et al. (2004). Borrowed from localisation physics, the IPR is I_k, the sum of the fourth powers of an eigenvector's components. Higher IPR means greater concentration; lower IPR means a more dispersed vector. The authors convert it into N_eff, the number of names carrying the portfolio, then ask whether those names reproduce the full eigenportfolio return. The resulting regressions report R\fd of 0.79 for the largest eigenvector of C, 0.82 for the second and 0.86 for v1.
The result resolves into sectors. Utilities generally occupy the long side of v1. The authors say this pattern changes during financial stress, and their charts bear that out. Materials dominate the short side from June 2008 to June 2011, alongside a smaller financials position. From 2020 to 2023, the short sleeve is almost entirely financials. Away from those periods, it is barely present.
Which norm governs N_eff?
Equation (2) spells out the equally weighted case. Each portfolio component is 1/p, giving I_k = 1/p\fb3, and equation (3) applies the cube root to recover N_eff = p. At the concentrated extreme, a single component equal to 1 produces I_k = 1 and N_eff = 1. The endpoints both behave as intended. For L1-normalised weights, the cube root is internally consistent.
An eigendecomposition, however, returns unit L2 vectors. Their equal components are 1/\fa1p, their IPR is 1/p, and the usual effective count is 1/I_k. The paper says only that "when normalised, the eigenvectors of this correlation matrix represent portfolios" and later uses eigenvectors to mean their normalised versions. The general-case norm remains unspecified.
The choice changes the result sharply. Put a near-uniform unit-L2 eigenvector into I_k^(-1/3), and the answer is 484^(1/3), about 7.85 effective names. The L1 interpretation gives 484. The gap is a factor of about 62 for the quantity that determines membership in the effective set. Figure 4 shows N_eff over time, although the printed text does not provide recoverable values. We therefore cannot identify the convention behind the figure.
Signs create a second problem for v1. Signed-L1 is the literal interpretation of equation (2), yet it becomes unstable for a long-short vector. As the signed loading sum approaches zero, all weights explode, I_k rises above 1 and N_eff falls below 1. Absolute-L1 matches equation (2) exactly when every component is positive and remains bounded. I would choose absolute-L1.
The ranking rule itself is clear. A footnote says components are selected by absolute weight and enter the effective portfolio with their original signs.
We did not find a convention for fixing the eigenvector's sign. Since an eigenvector is defined only up to a sign flip, placing utilities on the long side requires an orientation rule. Every implementation has to supply one.
The 0.82 remains in sample
The validation notation defines E(R_u1(j)) as a p-vector of individual-stock returns. The prose instead describes a regression between the full portfolio return and the effective-only portfolio return. Depending on the intended reading, the observations are either 76 window points or 76 by 484 stock-window points. Those interpretations give R\fd very different meanings. The printed index for the negative side runs from N\f9d_eff + 1 to p, which would include almost the whole cross-section. I read it as the bottom N\f9d_eff names.
Timing matters more. The effective set for window j is evaluated on returns from that same window j, making the reported fits contemporaneous by design. The authors make no forecasting claim. Readers should not treat 0.86 as an out-of-sample result.
The sensitivity analysis is substantial. Appendix A repeats the exercise with window lengths of 630, 756 and 882 days, corresponding to 36, 42 and 48 months of data, along with steps of 5, 20 and 40 days. Both halves of that appendix carry the analysis.
The directional-stability test offers less reassurance. Consecutive 504-day windows advanced by 60 days retain 444 common observations, or 88% of their data. A t-test on angular changes between heavily overlapping windows contains little independent information.
Utilities are plausible. The eigenvalue behavior is less obvious
Regulated cash flows and inelastic demand make utilities an unsurprising persistent long sleeve for a market-neutralised correlation mode. A sector desk could have guessed that result. The changing short side carries more interest, although the paper identifies it after the event. In the conclusions, the authors say they do not formally assess how this information could enter real-world portfolio construction. They leave out-of-sample testing for future work.
Two findings are less predictable. First, \f9b\f82 of C declines during the 2007 crisis and Covid, while \f9b'\f81 of the residual matrix increases. Even as those eigenvalues move in opposite directions in 2007 and Covid, u2 and v1 remain aligned above 0.95 across most of the sample. The 2007 episode coincides with a fall in \f81(u\f82, v\f81) below 0.85. Covid does not.
Second, removing the market mode generally raises the effective count. A prior in which the market explains everything would suggest the opposite. Appendix B also earns its space: the third and fourth eigenvectors display no persistent sector dominance, weakening the common practice of interpreting several leading eigenvectors as sector factors.
The authors disclose the uncomfortable parts. Their universe takes end-2024 S&P 500 membership and carries it backward to 2005, and they explicitly identify the survivorship issue. Citing Plerou et al. (2002), they argue that dominant modes are insensitive to the chosen firm subset. We did not find a test of that claim in the paper. GICS classifications are fixed as of Q1 2025. They also state plainly that real portfolio construction is outside their formal evaluation, an appropriate disclosure given their claim of improved transaction cost efficiency.
Our trading run
We could not reproduce the paper's universe because point-in-time S&P 500 membership back to 2005 is unavailable to us. We used an annual point-in-time universe of the top-500 US equities by capitalization, excluding ADRs. Our price history starts around 2010, which rules out replicating the 2005-2024 sample. The 504-day correlation window delays the tradable period further, leaving a run from 2020-01-01 to 2024-07-01.
It lost money.
From 2020-01-01 to 2024-07-01, total return was -35.93%, Sharpe -0.48, Sortino -0.54, Calmar -0.19, max drawdown -49.72% and volatility 16.80%. The paper gives no Sharpe, return or turnover for an N_eff strategy, leaving no equivalent author result for comparison.
Our implementation generated a daily signal and rebalanced monthly at the close. It used 504 daily log returns, the under-5% missing-day filter, single-index residualisation and signed-L1 normalisation of v1. We oriented v1 so the aggregate retained utilities loading was positive, set K_eff = floor(N_eff), and ranked names by absolute loading.
Positive-loading utilities formed the long sleeve. Negative-loading materials and financials formed the short sleeve, which opened only when lagged VIX was at or above 25 or SPY stood 10% or more below its trailing 63-day high. Gross exposure was 100/0 in ordinary periods and 50/50 under stress. Each position was limited to 10%, with leverage capped at 4.0. Every fill included commissions of $0.0040 per share, subject to a $1.00 minimum and a cap of 1% of trade value, before calculation of the metrics.
Three features of this construction limit what the loss can establish. Our orientation rule embeds the paper's conclusion by forcing utilities positive. We introduced the stress gate, whereas the paper identifies sector rotation retrospectively. The window includes one severe stress episode, so the short sleeve has almost no independent evidence supporting it.
Short borrow, financing on the leverage allowance and market impact are excluded. This result comes from one automated pass based on the paper's description. Its weaknesses bear first on our choices for normalisation and orientation, with much less force as evidence about the authors' work.
Evidence that would change the verdict
We have previously examined effective-asset counts that appear to be the result even though allocation drives the outcome (our note on crypto HRP variants). N_eff faces the same risk.
Two tests would resolve the issue. Calculate N_eff using absolute-L1, then using unit-L2 with the standard 1/I_k count, and establish whether both methods select the same sector sets. Next, choose the effective names in window j and score the eigenportfolio return in window j+1.
If the selected sets survive the first exercise and the returns survive the second, the method becomes a portfolio tool. For now, it remains a careful and useful description of the second eigenvector.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.