AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

The IPR eigenportfolio turns on one normalisation choice

Utilities remain on the long side, while the effective-name count changes by roughly 60x across conventions.

2026-09-10 · 8 min read · US equities

Reviewing: Bridging inverse participation ratio and portfolio theory · Laura Molero González, Roy Cerqueti, Juan Evangelista Trinidad Segovia et al. · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

IPR-Constrained Defensive Second-Eigenvector Utilities–Cyclicals Portfolio

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper

  • The exact historical S&P 500 constituent universe from 2005 onward cannot be reproduced because point-in-time index membership is unavailable; use a point-in-time top-500 US-equity capitalization universe instead. Our price history begins around 2010, so the paper's 2005-2024 sample cannot be replicated in full; evaluate the method over the available 2010-2024 period.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (11)

  • deviation left undescribed by the audit (invalidates: The paper's exact p=484 sample paths and J=76 windows; all date-specific Figure 1 through Figure 26 levels, overlaps, counts and sector distributions; direct replication of reported R-squared values)
  • deviation left undescribed by the audit (invalidates: The paper's crisis chronology covering 2007-2012; J=76; full-sample Figure 1 through Figure 26 paths and reported full-period regression R-squared values)
  • deviation left undescribed by the audit (invalidates: Exact J=76 and the paper's consecutive-window timing; direct comparison of the daily trading-signal path with Figures 1 through 8)
  • deviation left undescribed by the audit (invalidates: Any claim that strategy returns or turnover replicate the paper; the paper's qualitative defensive findings remain hypotheses rather than guaranteed trading results)

7 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total -35.9%Jul 2024
Sharpe
-0.48
Total Return
-35.9%
Max Drawdown
-49.7%
CAGR
-9.4%
Volatility
16.8%
Beta vs SPY
0.36
Trades
1,410

Equation (3) forces anyone implementing this paper to make a choice the text leaves unresolved.

First, the construction.

The portfolio Molero Gonz\filez and coauthors built

The authors take daily log returns for 484 S&P 500 names across 5,029 trading days, from 6 January 2005 to 21 November 2024. They estimate the correlation matrix over a rolling 504-day window, advancing it 60 days each time. The procedure produces 76 windows.

After normalisation, each eigenvector becomes a portfolio. The largest is the market mode. As a proxy, the authors use an equally weighted 1/p portfolio of the 484 names, with mean daily log return 0.0390%, standard deviation 1.2919%, skewness -0.6056 and a worst day of -13.58%.

The paper focuses on the second eigenvector. Earlier work by the same group found that the second-largest eigenvalue moves against the market mode under stress, making its portfolio a possible defensive overlay. The authors first remove the market through an OLS single-index regression of each stock on the market factor within each window. They then calculate the correlation matrix from the residuals. Its largest eigenvector, v1, serves as a stand-in for the second raw eigenvector.

The match is close. Pearson correlation stays above 0.95 through most of the sample. It falls between 0.90 and 0.95 from August 2010 to September 2012, and drops below 0.85 only during January to September 2007 and September 2022 to June 2023.

N_eff is the paper's new contribution. It is built from the Inverse Participation Ratio, used in finance since at least Utsugi et al. (2004). Borrowed from localisation physics, the IPR is I_k, the sum of the fourth powers of an eigenvector's components. Higher IPR means greater concentration; lower IPR means a more dispersed vector. The authors convert it into N_eff, the number of names carrying the portfolio, then ask whether those names reproduce the full eigenportfolio return. The resulting regressions report R\fd of 0.79 for the largest eigenvector of C, 0.82 for the second and 0.86 for v1.

The result resolves into sectors. Utilities generally occupy the long side of v1. The authors say this pattern changes during financial stress, and their charts bear that out. Materials dominate the short side from June 2008 to June 2011, alongside a smaller financials position. From 2020 to 2023, the short sleeve is almost entirely financials. Away from those periods, it is barely present.

Which norm governs N_eff?

Equation (2) spells out the equally weighted case. Each portfolio component is 1/p, giving I_k = 1/p\fb3, and equation (3) applies the cube root to recover N_eff = p. At the concentrated extreme, a single component equal to 1 produces I_k = 1 and N_eff = 1. The endpoints both behave as intended. For L1-normalised weights, the cube root is internally consistent.

An eigendecomposition, however, returns unit L2 vectors. Their equal components are 1/\fa1p, their IPR is 1/p, and the usual effective count is 1/I_k. The paper says only that "when normalised, the eigenvectors of this correlation matrix represent portfolios" and later uses eigenvectors to mean their normalised versions. The general-case norm remains unspecified.

The choice changes the result sharply. Put a near-uniform unit-L2 eigenvector into I_k^(-1/3), and the answer is 484^(1/3), about 7.85 effective names. The L1 interpretation gives 484. The gap is a factor of about 62 for the quantity that determines membership in the effective set. Figure 4 shows N_eff over time, although the printed text does not provide recoverable values. We therefore cannot identify the convention behind the figure.

Signs create a second problem for v1. Signed-L1 is the literal interpretation of equation (2), yet it becomes unstable for a long-short vector. As the signed loading sum approaches zero, all weights explode, I_k rises above 1 and N_eff falls below 1. Absolute-L1 matches equation (2) exactly when every component is positive and remains bounded. I would choose absolute-L1.

The ranking rule itself is clear. A footnote says components are selected by absolute weight and enter the effective portfolio with their original signs.

We did not find a convention for fixing the eigenvector's sign. Since an eigenvector is defined only up to a sign flip, placing utilities on the long side requires an orientation rule. Every implementation has to supply one.

The 0.82 remains in sample

The validation notation defines E(R_u1(j)) as a p-vector of individual-stock returns. The prose instead describes a regression between the full portfolio return and the effective-only portfolio return. Depending on the intended reading, the observations are either 76 window points or 76 by 484 stock-window points. Those interpretations give R\fd very different meanings. The printed index for the negative side runs from N\f9d_eff + 1 to p, which would include almost the whole cross-section. I read it as the bottom N\f9d_eff names.

Timing matters more. The effective set for window j is evaluated on returns from that same window j, making the reported fits contemporaneous by design. The authors make no forecasting claim. Readers should not treat 0.86 as an out-of-sample result.

The sensitivity analysis is substantial. Appendix A repeats the exercise with window lengths of 630, 756 and 882 days, corresponding to 36, 42 and 48 months of data, along with steps of 5, 20 and 40 days. Both halves of that appendix carry the analysis.

The directional-stability test offers less reassurance. Consecutive 504-day windows advanced by 60 days retain 444 common observations, or 88% of their data. A t-test on angular changes between heavily overlapping windows contains little independent information.

Utilities are plausible. The eigenvalue behavior is less obvious

Regulated cash flows and inelastic demand make utilities an unsurprising persistent long sleeve for a market-neutralised correlation mode. A sector desk could have guessed that result. The changing short side carries more interest, although the paper identifies it after the event. In the conclusions, the authors say they do not formally assess how this information could enter real-world portfolio construction. They leave out-of-sample testing for future work.

Two findings are less predictable. First, \f9b\f82 of C declines during the 2007 crisis and Covid, while \f9b'\f81 of the residual matrix increases. Even as those eigenvalues move in opposite directions in 2007 and Covid, u2 and v1 remain aligned above 0.95 across most of the sample. The 2007 episode coincides with a fall in \f81(u\f82, v\f81) below 0.85. Covid does not.

Second, removing the market mode generally raises the effective count. A prior in which the market explains everything would suggest the opposite. Appendix B also earns its space: the third and fourth eigenvectors display no persistent sector dominance, weakening the common practice of interpreting several leading eigenvectors as sector factors.

The authors disclose the uncomfortable parts. Their universe takes end-2024 S&P 500 membership and carries it backward to 2005, and they explicitly identify the survivorship issue. Citing Plerou et al. (2002), they argue that dominant modes are insensitive to the chosen firm subset. We did not find a test of that claim in the paper. GICS classifications are fixed as of Q1 2025. They also state plainly that real portfolio construction is outside their formal evaluation, an appropriate disclosure given their claim of improved transaction cost efficiency.

Our trading run

We could not reproduce the paper's universe because point-in-time S&P 500 membership back to 2005 is unavailable to us. We used an annual point-in-time universe of the top-500 US equities by capitalization, excluding ADRs. Our price history starts around 2010, which rules out replicating the 2005-2024 sample. The 504-day correlation window delays the tradable period further, leaving a run from 2020-01-01 to 2024-07-01.

It lost money.

From 2020-01-01 to 2024-07-01, total return was -35.93%, Sharpe -0.48, Sortino -0.54, Calmar -0.19, max drawdown -49.72% and volatility 16.80%. The paper gives no Sharpe, return or turnover for an N_eff strategy, leaving no equivalent author result for comparison.

Our implementation generated a daily signal and rebalanced monthly at the close. It used 504 daily log returns, the under-5% missing-day filter, single-index residualisation and signed-L1 normalisation of v1. We oriented v1 so the aggregate retained utilities loading was positive, set K_eff = floor(N_eff), and ranked names by absolute loading.

Positive-loading utilities formed the long sleeve. Negative-loading materials and financials formed the short sleeve, which opened only when lagged VIX was at or above 25 or SPY stood 10% or more below its trailing 63-day high. Gross exposure was 100/0 in ordinary periods and 50/50 under stress. Each position was limited to 10%, with leverage capped at 4.0. Every fill included commissions of $0.0040 per share, subject to a $1.00 minimum and a cap of 1% of trade value, before calculation of the metrics.

Three features of this construction limit what the loss can establish. Our orientation rule embeds the paper's conclusion by forcing utilities positive. We introduced the stress gate, whereas the paper identifies sector rotation retrospectively. The window includes one severe stress episode, so the short sleeve has almost no independent evidence supporting it.

Short borrow, financing on the leverage allowance and market impact are excluded. This result comes from one automated pass based on the paper's description. Its weaknesses bear first on our choices for normalisation and orientation, with much less force as evidence about the authors' work.

Evidence that would change the verdict

We have previously examined effective-asset counts that appear to be the result even though allocation drives the outcome (our note on crypto HRP variants). N_eff faces the same risk.

Two tests would resolve the issue. Calculate N_eff using absolute-L1, then using unit-L2 with the standard 1/I_k count, and establish whether both methods select the same sector sets. Next, choose the effective names in window j and score the eigenportfolio return in window j+1.

If the selected sets survive the first exercise and the returns survive the second, the method becomes a portfolio tool. For now, it remains a careful and useful description of the second eigenvector.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading date t:
  1. Select that calendar year's point-in-time top 500 non-ADR STOCK records
     by descending capitalization.
  2. Load 505 closes and compute 504 daily logarithmic returns.
  3. Remove stocks failing the strictly-less-than-5% missing-day rule.
     Use common valid dates without imputing returns; skip if n_observations < p.
  4. Standardize returns with ddof = 1 and form the Pearson correlation matrix C.
  5. Compute the second raw eigenvector u2 for diagnostics.
  6. Apply the implementation's Single Index market-removal procedure and compute
     the largest residual-correlation eigenvector v1.
  7. Normalize v1 by signed L1, then orient it so aggregate retained Utilities
     loading is positive; use the deterministic largest-component fallback if needed.
  8. Compute IPR I′1 = Σ(v1_i^4), N_eff = 1 / cubert(I′1), and K_eff = floor(N_eff).
  9. Rank stocks by descending |v1_i|, breaking ties by ascending symbol_key,
     and retain the first K_eff names while preserving their oriented signs.
 10. Long candidates: positive-loading Utilities.
     Stress-only short candidates: negative-loading Materials or Financials.
 11. Classify stress using information through t-1:
       VIX close >= 25, OR
       SPY / trailing 63-day SPY maximum - 1 <= -10%.
     If both inputs are unavailable, stress is unknown and no short sleeve opens.
 12. Store the daily signal and diagnostics, but generate orders only on monthly
     rebalance dates at the observed close using MOC execution.
 13. Outside stress, target 100% long and 0% short gross exposure.
     During stress, target 50% long and 50% short gross exposure when sleeves exist.
 14. Allocate each sleeve in proportion to |v1_i|, cap each position at 10%,
     and enforce maximum leverage of 4.0.
 15. Skip an order if its observed execution close is unavailable; never synthesize,
     carry forward, or impute an execution price.