AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

A t of 5.40 on a Premium With Mismatched Legs

Rotation prices a non-canonical variance premium; horizon-matched versions return |t| under 1.2

2026-08-29 · 12 min read · US equities and the SPY ETF, with VIX index levels used as a readable volatility-state input.

Reviewing: The Reconfiguration Premium: Co-movement Structure as an Unspanned Dimension of the Variance Risk Premium · Lucas Carvalho · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Persistent Correlation-Eigenspace Reconfiguration SPY-to-Cash Risk Overlay

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

Run on a different market than the paper

The paper estimates its correlation-structure measure on historical S&P 500 constituents, including delisted names. We cannot reproduce point-in-time S&P 500 membership, so use a yearly point-in-time top-500 US-equity universe selected from aiquant_screening_table.capitalization, trading SPY rather than an index itself. The mechanism—persistent rotation in the cross-sectional equity correlation eigenspace foreshadowing a broader volatility environment—can plausibly survive this closely related large-cap US-equity universe change, but the paper's reported magnitudes and significance do not transfer.

The paper's own figures describe its universe and do not carry over to ours.

This is not a replication of the paper

  • The original historical S&P 500 constituent panel, including point-in-time membership and delisted-name coverage from 1994, is not available; the test is limited to the platform's US-equity coverage from approximately 2010 and an annually defined top-cap proxy universe. The paper's Cboe COR1M, COR3M, and DSPX implied-correlation-index spanning tests are not available as specified, so those particular validation tests cannot be replicated. The proposed implementation tests a risk-overlay adaptation, not the paper's stated variance-risk-premium attribution result; the paper explicitly reports no timing alpha or crash-protection result.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (14)

  • deviation left undescribed by the audit (invalidates: The paper's reported VRP coupling coefficients, t statistics, confidence interval, R-squared values, premium decomposition, and variance-harvest attribution do not transfer to this SPY application.)
  • deviation left undescribed by the audit (invalidates: The reported beta of +0.195, Newey-West standard error 0.036, t statistic +5.40, confidence interval, R-squared, and mechanical-control premium coupling do not apply.)
  • deviation left undescribed by the audit (invalidates: The paper's capture rates, quartile payoff shares, loss frequencies, conditional losses, and short-variance Sharpe comparisons are not portfolio expectations for this strategy.)
  • deviation left undescribed by the audit (invalidates: The paper's linear short-variance timing Sharpe and paired t statistics do not apply to the SPY overlay.)

10 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 68.2%Jul 2024
Sharpe
0.73
Total Return
68.2%
Max Drawdown
-36.5%
CAGR
12.3%
Volatility
20.7%
Beta vs SPY
0.91
Trades
55

What the paper reports for its own strategy

  • Variance-harvest attribution (selling variance at VIX-squared against the paper's 12m equal-weighted realized leg, July 1998-November 2025, n=329, transaction costs, margin and VIX-squared convexity explicitly ignored, non-tradeable payoff): mean monthly payoff by real-time REC quartile +0.0040 (Q1), +0.0102 (Q2), +0.0092 (Q3), +0.0410 (Q4); capture rates 11.1%, 27.4%, 24.2%, 52.7%; P&L shares 6%, 16%, 14%, 64%
  • Risk by REC quartile over the same 329 months: probability of loss 0.29, 0.21, 0.26, 0.15; mean loss given loss -0.0265, -0.0164, -0.0181, -0.0102; worst month -0.099, -0.068, -0.064, -0.032
  • Double sort on VIX median then REC quartile: mean payoff +0.0032 (Q1) vs +0.0047 (Q4) in the low-VIX half (worst months -0.031 and -0.024); +0.0115 (Q1) vs +0.0551 (Q4) in the high-VIX half (worst months -0.099 and -0.024)
  • Timing overlay, pre-registered null: unconditional short-variance Sharpe 1.49 vs scaled (1 + 0.5 z_t) Sharpe 1.36, paired t = -1.85 at equal volatility over 329 months (against the overlay); Sharpes stated to be inflated by the non-tradeable payoff definition
  • Timing battery on the 293-month real-time-threshold sample: unconditional Sharpe 1.28; linear tilt Sharpe 1.22, equal-volatility t = -0.81; step tilt Sharpe 1.25, t = -0.94; top-quartile-only Sharpe 1.02, t = -2.23; step form by half-sample -0.97 (2001:07-2013:08) and -0.07 (2013:09-2025:11)
  • No crash protection: over the 329-month payoff sample the worst 5 percent of months carry a mean standardized REC reading of -0.47

Carvalho has found a real state variable and attached it to the wrong premium. The rotation index is unspanned by the level of the traded correlation surface, at most 6.7 percent, though its correlation with VIX is 0.316 and the slope of that surface does react. It forecasts the broad volatility environment two to three quarters out. The object it prices is a one-month implied variance against a twelve-month equal-weighted constituent realized leg. Match the legs and the coupling goes away. The mismatch is disclosed in the data section and again in the limitations, and the interesting question is what survives the disclosure.

The measurement

Every priced measure of co-movement in the literature is a function of correlation eigenvalues: index variance, average pairwise correlation, option-implied correlation, the absorption ratio. Carvalho's point is that the eigenvectors carry a separate piece of information, namely which firms co-move rather than how much, and that nothing prices their motion.

The construction. Estimate a twelve-month correlation matrix on S&P 500 constituents each month, work in the dual space because T = 12 is far smaller than the roughly 421-name effective cross-section, pull the loadings on modes two through four into stock space, orthonormalize to a three-dimensional subspace. Then measure the principal angles between this month's subspace and last month's, restricted to names present in both, and take the mean squared sine. Call it REC, the reconfiguration index. The market mode is dropped by construction, which is the choice that makes the whole thing work. The paper cites a market-inclusive rotation measure that correlates 0.42 with the irreversibility of a regime-aware ranking chain and near zero with its market-neutral counterpart, which is why that version tracks stress (Halperin, 2026b).

The reading is exact rather than metaphorical. One minus REC is the share of last month's co-movement structure retained. Sample mean 0.194, standard deviation 0.119, range 0.005 to 0.526 with the maximum in July 2002. A typical month rewrites a fifth of the classification and carries four-fifths forward, about a 26-degree tilt of the subdominant frame.

Data: monthly total returns, June 1994 to December 2025, including delisted names, filtered to a 379 x 430 panel of 159,742 observations. That yields 368 windows, 367 rotation readings, 365 regression observations.

The headline. Regress the log premium (one-month VIX squared minus twelve-month equal-weighted constituent realized variance) on the standardized three-month average of REC, with realized volatility and average pairwise correlation as controls, Newey-West at twelve lags. Beta +0.195, standard error 0.036, t = +5.40, R-squared 0.504. On the raw unsmoothed index, t = +4.02. Both level controls enter negatively and significantly, so the coupling is identified against volatility rather than through it.

What it is supposed to buy you: the premium is prepayment. On impact, the whole association sits in implied variance (t = +4.57) and none in realized (+0.62). Forward realized variance builds to a peak coefficient of +0.155 at eight months (t = 3.65), and the premium itself goes from +5.12 at h = 0 to +0.22 at three months and troughs at -2.68 at eight. Money in advance for turbulence that then shows up.

The two premiums

The paper's dependent variable mismatches its legs in horizon (one month against twelve) and in weighting (cap-weighted index implied against equal-weighted constituent realized). Carvalho states this as a scope condition in the data section rather than a footnote. He also reruns a companion battery to bound it. The same smoothed index gives |t| <= 1.2 against horizon-matched forward variance-swap payoffs at one, three and six months (-0.08, +0.74, +1.19), and |t| <= 1.9 against forward one-month cap-weighted index realized variance at every horizon tested.

Those two lines do most of the work in this review. The 5.40 and the -0.08 measure different objects. REC prices the spread between short-dated index fear and the slow, broad volatility environment of the persistent large-cap cross-section. It does not price the one-month variance swap.

The author's defence of that is better than it first looks, and it deserves to be quoted whole rather than clipped. The paper says the non-porting result "is not a weakness of the result but a condition of its consistency with the no-alpha boundary": a state variable that forecast the variance-swap payoff would be a trading signal, and this one forecasts the environment against which the premium is written. Internally that is coherent. The headline t of 5.40 still attaches to a quantity no desk has a book in, and the paper says as much twice: equation (2) "is not the canonical variance risk premium", and the harvest payoff "is not the payoff to a tradeable variance swap". The abstract reports the coupling to the aggregate variance risk premium at t = 5.40, with no scope condition attached to the word aggregate. The disclosure sits in Sections 2, 6.4 and 8.4 instead.

And the convention matters too. R-squared is 0.504 in logs against 0.149 in levels. October 2008 through December 2009 accounts for 46 percent of the total squared deviation of the levels premium about its mean. REC's own largest readings are March 2022, October 2018, July 2010, April 2000 and May 2020. Those months sit against levels premia between -0.015 and +0.048, versus a crisis maximum of +0.299. Rotation prices ordinary variation in the premium, not its extremes. Carvalho says this outright and links it to the no-crash result.

What the estimator survives

I expected this measure to be a laundered version of correlation intensity, and the paper anticipates that. Spectral gaps enter with the negative signs Davis-Kahan predicts (t = -3.14 upper, -2.15 lower). The full mechanical model explains only R-squared 0.121 of the index's variance. The pricing coupling holds at t = +4.92 with the whole battery as controls. Lower-gap quintile couplings run +2.02, +3.25, +3.47, +3.26, +0.97, no monotone pattern, and the calmest variance-share quintile still gives +3.96. The sorts also show where it is absent: +0.77 in the middle correlation quintile, +1.67 and +1.76 in variance-share quintiles two and three. The paper prints those.

Orthogonality to the level gauges holds. Across the correlations of the index with level gauges, no cell exceeds 0.32; the largest is the smoothed index against VIX at 0.316. Realized volatility is -0.031 raw, average pairwise correlation -0.056. Projecting REC on COR3M (Cboe's three-month implied correlation index) and COR1M recovers R-squared 0.059 on the raw index and 0.050 on the smoothed, over 240 months from 2006. Adding DSPX (the S&P 500 dispersion index) gives 0.052 raw and 0.067 smoothed over 139 months from 2014. Between 5.0 and 6.7 percent spanned, depending on smoothing and sample. Meanwhile the shape of that same surface reacts: the COR3M minus COR1M slope loads on smoothed REC at t = -3.46, beta -0.910 index points per standard deviation against a mean slope of +3.409. The level instruments do not carry the state variable, and the paper gives a Rayleigh-quotient argument for why: for the average-correlation functional the median squared overlap with the market mode is 0.750 against body-mode overlaps of 0.013, 0.004 and 0.002.

Smoothing is not load-bearing. t(REC) across averaging windows of 1 to 24 months runs 4.02, 4.00, 5.40, 5.46, 5.34, 5.03, 4.06, 3.69, 3.28, 3.09, 3.64, and the adopted K = 3 sits below the K = 4 and K = 5 maximum. Panel-filter sensitivity spans t of +2.44 to +6.11 across seven threshold pairs, again with the adopted pair not the maximum. The K = 3 subspace, the deflated-edge discard rule and the three-month window remain in-sample conventions. Carvalho concedes that part of the K = 3 to K = 5 decline may be subspace-dimension mechanics rather than bulk dilution.

One place the inference genuinely wobbles. A block-pairs bootstrap is the natural first choice for overlapping regressions. It raises 95 percent critical values to 3.63-5.59 and kills every horizon. The paper reports the failed design and argues that blocks preserve the lead-lag structure that is the alternative, so the null absorbs the effect. I find that argument persuasive. The forward result still stands on the i.i.d.-pairs null, and the h = 10 verdict flips between AR(2) and AR(4).

What rotates, in desk language

Mode two is labelled utilities in 63.6 percent of the 368 windows and energy in 17.7 percent; mode three is diffuse in 52.2 percent, mode four in 71.2 percent. One axis does most of the turning: median principal angles are 5.9, 11.8 and 42.8 degrees, and in March 2022 they were 7, 10 and 71. Two axes hold, one turns.

So the priced thing is drift in which names load on a duration-and-defensives axis, with a commodity axis as the frequent alternative. Discrete label switching does not absorb it: a switch-intensity series correlates +0.187 with the raw index and leaves it at t = +4.89 against the switch series' +1.93. Carvalho's own framing is loading drift, identified by elimination rather than direct measurement, and he flags that as a limitation. Translated: your hedge ratios and your sector buckets are being re-estimated by the market underneath you, and the pace of that re-estimation over a quarter is what carries the premium. The monthly innovation carries nothing (t = +0.30) while the three-month average carries everything (t = +4.92). The premium pays for pace, and one-month rotation drives out the twelve-month version, 4.76 against 1.38.

One more number that stuck with me. Subspaces 24 months apart still retain 0.0664 alignment against a chance level of 3/421 = 0.0071, verified by 3,000 simulated pairs. Nine times chance, and the profile is flat from twelve months out. A permanent core plus a transient with roughly twelve-month memory, and the forecast dies at h = 11-12, exactly where the subspace has decorrelated.

Where the money isn't

The attribution table is the part a variance seller will read first, and it is the part Carvalho most insistently labels as not a backtest. Selling variance at VIX squared against the paper's own realized leg, July 1998 to November 2025, 329 months, with transaction costs, margin and the VIX-squared convexity correction explicitly ignored and the payoff non-tradeable. Mean monthly payoff by real-time REC quartile: +0.0040, +0.0102, +0.0092, +0.0410. Capture rates 11.1, 27.4, 24.2, 52.7 percent. Sixty-four percent of 27 years of premium sits in the top rotation quartile, and that quartile is also the safest: loss probability 0.15 against 0.29, worst month -0.032 against -0.099. Compensation up, risk down, across the same sort. Whatever that is, it is not a risk exposure.

And it converts into nothing. The pre-registered linear overlay gives Sharpe 1.36 against 1.49 unconditional, paired t = -1.85 at equal volatility, against the overlay. On the 293-month real-time-threshold sample: unconditional 1.28, linear tilt 1.22 (t = -0.81), step tilt 1.25 (-0.94), top-quartile-only 1.02 (-2.23). Split in half, the step form gives -0.97 on 2001:07-2013:08 and -0.07 on 2013:09-2025:11. Over the 329-month payoff sample the worst 5 percent of months carry a mean standardized REC reading of -0.47. The paper states that both the 1.49 and the 1.36 are inflated by the payoff definition and should be read only as a relative comparison.

Can our adaptation settle anything?

What we could not do first. The original point-in-time S&P 500 constituent panel with delisted-name coverage from 1994 is not available to us, so we cannot rebuild the paper's 430-name survivor-tilted universe. Our coverage starts around 2010. The Cboe COR1M, COR3M and DSPX spanning tests are not available as specified, so those validation results are untested here. Most importantly, we did not test the paper's claim at all: we built a risk-overlay adaptation, while the paper's stated result is a variance-risk-premium attribution and it explicitly reports no timing alpha and no crash protection. We also traded SPY against cash rather than an index book or a variance swap.

What we ran instead. REC on a yearly point-in-time top-500 US-equity universe selected by capitalization, monthly returns, twelve-month windows, modes two to four, three-month smoothing, expanding standardization with a 36-month burn-in clipped to [-2, 2]. Then a gate. We recursively forecast one-month SPY realized variance with VIX plus lagged realized variance, and again with standardized persistent REC added. The two forecasts are scored against each other with QLIKE, a variance-forecast loss function. The overlay only acts once the augmented model's mean QLIKE beats the baseline on at least twelve completed forecast pairs. When active, SPY weight is clip(1 - 0.5 x max(z, 0), 0, 1) against cash at 0 percent. One basis point per one-way notional plus platform commissions, executed at the first SPY close after month-end. Backtest 2020-01 to 2024-07, 55 trades.

Our figures, from our run: Sharpe 0.73, max drawdown -36.5 percent, beta 0.91, annualized volatility 20.70 percent. The paper's own reported Sharpes are 1.49 unconditional and 1.36 for its scaled overlay over 329 months. The two sets measure different instruments: theirs is an uncosted, non-tradeable short-variance carry series on its own realized leg, ours is a long-only SPY-versus-cash equity overlay net of costs. The difference in level tells you about the change of traded object, the sample and the exposure rule. It tells you nothing about whether REC prices the paper's premium.

The beta is the number that explains our result.

At beta 0.91 and 20.70 percent annualized volatility, the overlay carried essentially full equity risk into a -36.5 percent drawdown through the 2020 crash. Part of that is structural in our own design. The twelve-month lookback, the three-month smoothing, the 36-month burn-in and the twelve-pair QLIKE gate consume most of the 2020-01 to 2024-07 window. So the overlay is likely sitting at 100 percent SPY for a large share of the run. That is a choice of ours, and it is the first thing I would change. Our forecast target is also cap-weighted SPY variance, and the paper reports |t| <= 1.9 for its own index against exactly that object, so a gated-off or ineffective overlay is what its evidence predicts for our construction.

One automated pass built from a paper's description is evidence about our implementation before it is evidence about anything else. What I will say is that the qualitative shape agrees with what Carvalho pre-registered: full retained beta and an unshortened drawdown are what "no timing alpha" and "no crash protection" look like once you wrap the signal in something tradeable.

What would change my mind

A direct measurement of loading drift on a churning cap-weighted panel. The index as built restricts each comparison to names present in both windows, so entering and exiting names contribute no rotation, and the identification of loading drift proceeds by elimination rather than measurement. Run the same construction where membership actually turns over, and measure the drift in loadings directly, and the mechanism claim stops resting on two ruled-out channels. The horizon-matched battery has already been run and returned |t| <= 1.2 against forward variance-swap payoffs at one, three and six months, so that door is closed. The daily-frequency version is closed too: the paper states that daily-sampled subspace geometry is a distinct object and that the companion paper finds it unpriced. As it stands the paper has established something worth knowing. The eigenvector orientation of the equity cross-section is priced information that the level of the traded correlation surface spans at most 6.7 percent of, and Carvalho has been unusually careful about the boundary of that claim. We have written before about how much of a covariance-model edge survives contact with a tradeable wrapper, in our note on MINGLE's covariance swap.

Read it as a correlation-regime diagnostic. Do not size off it.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each month-end t:
  1. Select the point-in-time top 500 stocks by capitalization.
  2. Form monthly adjusted-close returns and retain only stocks with complete
     observations within each trailing 12-month estimation window.
  3. Standardize each retained return column to zero mean and unit variance.
  4. Compute G_t = Z_t Z_t' / N_t and its ordered eigendecomposition.
  5. Map dual modes 2, 3, and 4 into stock space and QR-orthonormalize them.
  6. For windows t-1 and t, restrict both bases to common stocks and separately
     re-orthonormalize them.
  7. Compute principal angles and raw REC:
       REC_t = mean_i[sin(theta_i)^2]
             = 1 - ||A_t' B_t||_F^2 / 3.
  8. Smooth REC with its trailing three-month mean.
  9. For the portfolio signal, standardize smoothed REC using an expanding window,
     require a 36-month burn-in, and clip the z-score to [-2, 2].
 10. Recursively estimate one-month-ahead SPY realized-variance models:
       baseline = VIX + lagged SPY realized variance;
       augmented = baseline + standardized persistent REC.
 11. Evaluate only forecast pairs whose realized-variance outcomes are complete by t.
     Activate after at least 12 pairs only if mean QLIKE(augmented) is strictly below
     mean QLIKE(baseline).
 12. Set target exposure:
       if inactive: SPY weight = 1;
       if active:   SPY weight = clip(1 - 0.5 * max(z_REC_t, 0), 0, 1);
       cash weight = 1 - SPY weight.
 13. Execute at the first available SPY close after month-end and apply the target
     only to subsequent returns. Deduct one basis point per unit of one-way SPY
     notional traded.