AQAI QuantAI research lab for systematic strategies

Our review of the paperwe backtested it

The style timer earns its keep on drawdowns

Continuous macro signals cut the 2022 loss to a third of buy-and-hold; static growth still wins on CAGR.

2026-08-06 · 5 min read · US ETFs

Read the paper on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Continuous Smooth Macro-Market Growth vs Defensive ETF Style Timing

Backtest period 2020-01-01 to 2025-10-08 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper (3)

  • Exact VIX index history is not listed as an available data source; implementation should use an explicitly documented substitute such as SPY EOD option implied volatility where available or SPY realized volatility from daily_prices. This tests the same volatility-stress-relief mechanism but is not an exact replication of the paper's VIX-based signal.
  • Exact TNX index data is not listed, but the 10-year Treasury yield can be proxied using the available FRED DGS10 macro_indicators series.
  • Kenneth French daily factor data used for the paper's attribution is not in the provided catalog; the trading policy can be backtested without that attribution, but exact FF5+momentum regression diagnostics would require an external factor dataset.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (8)

  • eq 2 FF5+MOM attribution regression (invalidates: Full-sample FF5+MOM attribution for G-D and its reported betas, alpha, Newey-West t-statistic, and adjusted R^2)
  • eq 5 high-VIX state vh_t = z(VIXPercentile_{756,t}) (invalidates: Selected Smooth Score 2017-06-28 to 2026-05-15 performance; OOS validation performance; post-2022 validation performance; score sorting diagnostic)
  • eq 6 VIX relief vr_t = -z(DeltaVIX_{21,t}) (invalidates: Selected Smooth Score 2017-06-28 to 2026-05-15 performance; OOS validation performance; post-2022 validation performance; score sorting diagnostic)
  • eqs 25-30 expanded local grid (invalidates: Expanded local grid top-five ranking; statements about selected configuration not being the highest raw CAGR in the local grid)

4 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 150.9%Oct 2025
Sharpe
0.92
Total Return
150.9%
Max Drawdown
-51.1%
CAGR
17.3%
Volatility
28.6%
Trades
13,380

What the paper reports for its own strategy

  • Selected smooth-score policy, 2017-06-28 to 2026-05-15, 10bp cost: 19.24% CAGR, Sharpe 1.01, Sortino 1.22, vol 19.29%, max drawdown -31.63%, turnover 469.67%
  • Strict incremental Old+Credit overlay, 2017-06-28 to 2026-05-15, 10bp cost: 19.80% CAGR, Sharpe 1.04, max drawdown -31.92%, turnover 410.23%
  • OOS walk-forward expanding, 2018-06-28 to 2026-05-15, 10bp cost: 18.64% CAGR, Sharpe 0.96, max drawdown -32.93%
  • OOS walk-forward rolling, 2018-06-28 to 2026-05-15, 10bp: 18.02% CAGR, Sharpe 0.93; fixed parameter: 17.86% CAGR, Sharpe 0.93
  • Post-2022 WF expanding, 2022-01-03 to 2026-05-15, 10bp: 15.30% CAGR, Sharpe 0.90, max drawdown -19.89%
  • Old+Credit rolling OOS, 2018-06-28 to 2026-05-15, 10bp: 20.24% CAGR, Sharpe 1.03, max DD -32.36%

The paper's own comparison table settles the main question before you get to the machinery. Over June 28 2017 to May 15 2026, net of 10bp one-way costs, the timing policy returned 19.24% CAGR at Sharpe 1.01. Holding the growth basket outright returned 21.34% CAGR at Sharpe 0.94. So the strategy buys you 0.07 of Sharpe and about 2.7 points of drawdown relief (-31.63% versus -34.35%), and it costs you a little over 2 points of CAGR a year. Whether that trade is worth making turns entirely on whether you can sit through a 34% drawdown in a single growth sleeve. Most books cannot. The honest case for this paper is that narrow, narrower than a general claim that macro signals improve allocation.

What is actually being timed

G is an equal-weight basket of QQQ, XLK, VGT, SPYG, VUG. D is SCHD, VYM, VTV, FDVV, COWZ. The relative bet G minus D is a plain style exposure: market beta 0.273, HML -0.552, momentum 0.117, and annualized alpha of 1.95% with a Newey-West t of 0.81. That t settles the question of alpha. There is none to harvest here, and the author says as much. Every dollar the strategy makes comes from timing a known long-growth, short-value tilt. Read the whole thing as duration management on equities, because that is what the factor loadings say it is.

The score replaces regime labels

The design contribution is that the allocation is smooth. Instead of labelling days "panic" or "overheated" and flipping between states, the policy builds a continuous target growth weight from four inputs (rate relief, SPY drawdown depth, a volatility-stress term, and a growth-crowding penalty), maps it through a tanh, and eases into it with an EWMA rather than switching. The score-sorting diagnostic gives it some credit. The top quintile of the smooth score beat the bottom by 6.48% of forward G minus D return, against 3.82% for a rate-only score. Bigger tilts helped monotonically: a 20% maximum tilt gave 17.89% CAGR at Sharpe 0.95, a 50% tilt gave 19.02% at 0.99, with turnover climbing from 186% to 466%. The 50% version is the one carried forward.

Turnover above 400% a year, on a strategy whose entire edge is 0.07 of Sharpe, is the number I would stare at. The cost sensitivity holds up in the paper's own stress test (the selected policy stays above the 50/50 benchmark even at 20bp), but the margin is thin enough that your real fill quality decides whether this survives.

The validation, and the era it sits in

The out-of-sample work is real: walk-forward expanding and rolling, a fixed-parameter run, and a post-2022 check. The strongest single result is post-2022. Expanding walk-forward returned 15.30% CAGR against 100% growth's 15.45%, while cutting max drawdown from -33.92% to -19.89%. Nearly the same return, roughly half the drawdown. The yearly split shows the mechanism cleanly. In 2022 the growth basket lost 30.61% and the policy lost 9.53%. In 2023 growth rebounded 48.31% and the policy captured 29.38%. It clips both tails. Whether clipping helps depends on the regime, and the full 2017 to 2026 sample is a growth-friendly one. Run this exact strategy across a decade where value led and it would look different, and the paper cannot show you that decade.

Credit is a throttle

The bond/credit extension is the tidier finding. Added as a replacement signal, credit underperformed (17.73% CAGR, Sharpe 0.94). Added as a strict overlay on the existing score, it improved things: 19.80% CAGR, Sharpe 1.04, with turnover falling from 469.67% to 410.23%. The credit term earns its place mostly by trimming average growth exposure and turnover, not by adding return. One HAC t of 2.36 on a rate-relief-by-credit-stress interaction is the stated reason it was carried in. Thin, but the overlay does the operationally sensible thing.

What we could run, and what we could not

We could not reproduce the signal exactly. The volatility-stress block is built from VIX percentiles in the paper; VIX index history was not available to us, so we substituted a 21-day annualized realized SPY volatility series for that entire block. Same mechanism (volatility stress and relief), measured a slower way. We proxied the 10-year yield with FRED's DGS10. The Kenneth French daily factors behind the attribution were not in our catalog, so we did not attempt the FF5-plus-momentum regressions; we backtested the trading policy only. And we ran a shorter window, 2020-01-01 to 2025-10-08, opening straight onto the COVID crash with almost no pre-sample buffer for the expanding z-score.

With those substitutions, our run made 17.34% CAGR at Sharpe 0.92. Close to the paper's 19.24% and 1.01 on return and Sharpe, and since these come from a different volatility signal over a different period, read the gap as a difference in construction rather than a failed replication. The risk side is where we diverge hard. Our realized volatility was 28.57% against the paper's 19.29%, and our max drawdown was -51.11% against -31.63%. For a strategy sold on drawdown control, a 51% drawdown is close to a verdict against our own build. Two of our choices most likely explain the direction. The realized-vol proxy reacts to a fast repricing later than a VIX percentile would, so the stress-relief and crowding terms fire late and the allocator de-risks after the damage is done. And starting in January 2020 drops the crash into the first weeks, when the expanding z-score has too little history to normalize, so early tilts are noisy and mistimed. We cannot fully account for the size of the drawdown gap with what we can see. The near-doubling is evidence about our volatility proxy and our window first, not about the author's signal.

If you already run a growth sleeve and your binding constraint is surviving a 2022-style repricing, a smooth style timer that gives up two points of CAGR to halve the crash is worth costing out at your real turnover. If you came looking for alpha, the t of 0.81 answered you on the first page.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading day t:
  Build equal-weight growth basket G from QQQ, XLK, VGT, SPYG, VUG.
  Build equal-weight defensive basket D from SCHD, VYM, VTV, FDVV, COWZ.

  Using point-in-time data available for signal date t:
    Compute 21-day change in DGS10 and set rate_relief r_t = -z(DeltaTNX_21).
    Compute SPY drawdown depth d_t = -z(SPY_drawdown).
    Compute 21-day annualized realized SPY volatility.
    Compute 756-day volatility percentile and 21-day volatility change.
    Compute VIX-style stress variables from the single realized-volatility proxy.
    Compute 126-day trailing return of G minus D and standardize it.

  CoreScore_t = alpha * r_t + (1 - alpha) * d_t.
  StressScore_t = 0.5 * z(r_t * high_vol_state) + 0.5 * z(high_vol_state * vol_relief).
  CrowdedScore_t = 0.5 * z(growth_extension * low_vol_state)
                 + 0.5 * z(growth_extension * low_vol_state * rate_quiet).
  RawScore_t = CoreScore_t + lambda_s * StressScore_t - lambda_c * CrowdedScore_t.
  ScoreHat_t = expanding_zscore(RawScore_t).

  target_growth_weight = 0.5 + MaxTilt * tanh(ScoreHat_t / tau_w).
  actual_growth_weight = (1 - eta) * prior_growth_weight + eta * target_growth_weight.
  defensive_weight = 1 - actual_growth_weight.

  At the next trading-day close, allocate:
    each growth ETF weight = actual_growth_weight / 5
    each defensive ETF weight = defensive_weight / 5
  Apply transaction costs proportional to absolute change in growth allocation.