AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Why the 11-day covariance carries the result

Devanathan, Tzikas and Boyd raise 60/40 from 0.56 to 1.08 Sharpe before their end-to-end bootstrap removes the edge

2026-09-09 · 10 min read · US ETFs (SPY, AGG, GLD) plus cash

Reviewing: Simple Dynamic Stock/Bond/Gold Portfolios · Nikhil Devanathan, Alexandros E. Tzikas and Stephen P. Boyd · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Constrained SPY-AGG-GLD Momentum and Macro Markowitz with Cash and Volatility-Controlled Benchmarks

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (9)

  • All FRED observations must be point-in-time and respect publication availability. (invalidates: Exact paper Markowitz return, Sharpe, drawdown, allocation, and statistical-significance results; exact historical cash and excess-Sharpe replication)
  • Paper uses DFF; task requests FEDFUNDS. (invalidates: Exact paper cash return; exact geometric federal-funds-excess Sharpe; all quoted paper performance values)
  • Paper ridge uses 42 features. (invalidates: Paper Markowitz Table 1 values; Markowitz subperiod Sharpes; Markowitz average allocations; Markowitz frontier dominance; Markowitz bootstrap intervals and paired Sharpe differences; Markowitz lag-sensitivity values)
  • Paper uses its 2006-01-01 to 2026-01-01 simulation period. (invalidates: All paper full-period and subperiod numerical results; 2008 return comparisons; paper bootstrap confidence intervals)

5 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 18.9%Jul 2024
Sharpe
0.43
Total Return
18.9%
Max Drawdown
-16.4%
CAGR
3.9%
Volatility
9.8%
Beta vs SPY
0.28
Trades
120

What the paper reports for its own strategy

  • Markowitz: 11.6% annualized return, 9.0% annualized volatility, Sharpe 1.08, max DD 18.1%, mean DD 3.0%, turnover 284.4%, 2006-2026, net of 5bp trading cost
  • Simple Markowitz: 10.4% return, 9.3% volatility, Sharpe 0.91, max DD 15.8%, turnover 278.9%, same period and costs
  • 50/30/20 volatility-controlled (7% target): 7.8% return, 7.3% volatility, Sharpe 0.82, max DD 15.9%, turnover 79.0%
  • 60/40 volatility-controlled (7% target): 7.1% return, 7.4% volatility, Sharpe 0.71, max DD 16.9%, turnover 85.0%
  • Markowitz five-year subperiod Sharpes 1.34 / 0.86 / 1.16 / 0.97 (2006-2011, 2011-2016, 2016-2021, 2021-2026)
  • Markowitz CPI-adjusted real return 8.8%, real Sharpe 0.99 (excess of realized inflation)

Lag the return forecast by a full month and Sharpe slips from 1.08 to 1.00. Lag both the return forecast and covariance estimate by that month, as the paper's Table 12 does, and Sharpe drops to 0.73. Those appendix D results contain the paper's real economic result. The abstract leads with the volatility overlay, which the authors say improves every risk-adjusted and drawdown metric. Appendix D assigns more importance to covariance than alpha. The choice of the regression portfolio as the headline book therefore needs explaining.

Three ETFs, cash and a 7% cap

The investable set contains SPY, AGG, GLD and cash accruing at the effective federal funds rate. The portfolios are long only, use no leverage or derivatives, and rebalance monthly from public data. Open-source code reproduces every result. The test runs from January 2006 to January 2026, twenty years, and charges a flat 5bp bid-ask spread on every trade. The optimizer anticipates that charge in its objective. The authors frame the contribution as empirical rather than methodological, accompanied by a reproducible implementation that a self-directed investor could operate.

Six portfolios make up the comparison. Two are fixed-weight benchmarks rebalanced annually: conventional 60/40 and a 50/30/20 stock/bond/gold allocation motivated by rising stock-bond correlation. Another two apply volatility control to those benchmarks. Every month, the allocation moves toward cash until estimated ex-ante volatility reaches 7% annualized. That estimate comes from the benchmark's own daily returns over an 11-trading-day trailing window.

The remaining two portfolios optimize weights. Each maximizes the same monthly objective: forecast return, plus the cash return earned by the unallocated share, minus an L1 turnover cost. Two constraints do the heavy lifting. Estimated portfolio volatility has a hard 7% ceiling, while an L1 trust region holds relative weights within distance 1 of (0.5, 0.3, 0.2). The portfolio is long-only, and weights sum to at most one. Covariance comes from the annualized 11-day trailing sample matrix.

Only the return forecast changes between them. Simple Markowitz takes an EWMA of past daily returns with a 252-day half-life, then scales it by 21. The full Markowitz portfolio predicts 100-day forward returns through a ridge regression using 42 features, lambda 10, at most 512 training observations and 252-day half-life weights. Inputs include trailing returns and volumes, Treasury yields, breakevens, smoothed VIX, the dollar index and eight Fama-French factors. Its final training target is dated t minus 100, preserving causality.

After the 5bp spread, Markowitz returns 11.6% with 9.0% volatility, Sharpe 1.08, an 18.1% maximum drawdown and 284.4% turnover. Simple Markowitz records 10.4%, 9.3% and Sharpe 0.91. The volatility-controlled 50/30/20 portfolio produces 7.8%, 7.3% and 0.82. Conventional 60/40 returns 8.1% at 11.3% volatility, with Sharpe 0.56 and a 33.7% drawdown. Both optimization portfolios stayed positive through 2008, returning +4.0% and +2.6%.

Where does the alpha come from?

The authors rule out the simplest account themselves. A static portfolio holding Markowitz's average weights, 41.5% SPY, 20.1% AGG, 23.3% GLD and 15.1% cash, returns 8.5% at 9.1% volatility when rebalanced annually. Its Sharpe is 0.73. Risk remains similar while risk-adjusted return falls to two thirds of the dynamic portfolio's figure. Variation supplies the advantage. In the authors' words, the benefit comes from "varying the cash holdings (as well as the mix of stocks, bonds, and gold) strategically."

Which variation matters?

The lag test separates the components. Delaying alpha by 21 trading days preserves 1.00 of the original 1.08. Delaying alpha and covariance together by 21 days leaves 0.73, below the 0.82 from the basic volatility-controlled 50/30/20 portfolio. A five-day delay to both reduces Sharpe to 0.81. The authors reach the same reading: a reactive covariance forecast "appears to be substantially more important than a reactive return forecast."

Appendix F makes the point harder to avoid. Seven standard risk-based rules use the same three assets, costs and constraints. Their volatility-controlled Sharpes range from 0.60 to 0.85. Each rule reaches its highest Sharpe with the 11-day covariance window selected for the authors' portfolio, and each deteriorates at 21, 63, 126 and 252 days. Global minimum variance declines from 0.60 to 0.37 with the 252-day window. The authors answer this evidence in the same appendix: all seven rules share the 11-day covariance estimate, yet none exceeds 0.85. They conclude that "sizing positions by risk is evidently not enough on its own; the lift comes from combining a return forecast with a hard limit on estimated risk."

Estimating a three by three covariance matrix from eleven daily observations is aggressive. The authors acknowledge this, describing the window as "short by the standards of volatility estimation" and the result as "a noisy estimate". Its appeal is speed, and they selected it "after a modest search over various values". The rationale makes sense. During 2008 and March 2020, a useful risk estimate needed to react within days. Yet the full sample was already visible when the window was chosen. The paper tabulates a covariance sweep from 11 to 252 days for the seven comparison rules, all peaking at 11 days. Appendix E gives only a qualitative account for the authors' own portfolio, observing that the covariance estimate window "does affect the performance, though smoothly."

The alpha model still deserves credit. Black-Litterman reaches just 0.77 with volatility control when given exactly the same regression forecast. Across 90 specifications swept in-sample, its best Sharpe is 0.95, short of 1.08. Those winning specifications set view confidence near zero and effectively disable the equilibrium prior. Here, the shrinkage Black-Litterman was designed to supply reduced performance.

The end-to-end bootstrap removes the margin

The paper reports two bootstraps.

First, it resamples portfolio returns using 21-day mean blocks, 10,000 replications and paired observations across portfolios. Markowitz beats 60/40 by +0.52 Sharpe, with interval [+0.07, +0.99] and P(difference at most zero) = 0.011. The result clears 5%. Against 50/30/20, the difference is +0.38 and the interval is [-0.01, +0.79]. Against simple Markowitz, it is +0.17 with [-0.06, +0.40] and P = 0.068. The headline portfolio therefore beats the classic benchmark significantly, while no other comparison reaches significance. The authors say this directly. They also explain that simple Markowitz and the volatility-controlled portfolios are their own constructions, so those tests help choose among their portfolios rather than establish superiority over the status quo. That framing is fair.

The second bootstrap carries more weight. It uses a mean block length of 300 days and 100 replications. Raw inputs are resampled against a common index, after which the features are rebuilt, the ridge is refitted and all six portfolios are rerun. The result is blunt: "Every interval for the Markowitz Sharpe advantage contains zero, and the estimated probability of a nonpositive difference ranges from 0.28 to 0.51." The mean difference against 50/30/20 VC is -0.01. Against 60/40, it is +0.17 with interval [-0.23, +0.64] and P = 0.280. The authors describe this as a conservative check based on 100 replications, rather than a precise tail calculation. They are right, and they publish the result anyway, which is uncommon.

The sample itself raises another concern. Its strongest five-year block includes 2008 and posts Sharpe 1.34. Removing 2008 and 2009 almost preserves the ranking, at 1.11, 0.95, 0.93, 0.83, 0.87 and 0.74, though Markowitz's lead over 50/30/20 VC narrows from 0.26 to 0.18. During the 2011 to 2016 subperiod, the flagship portfolio ranks behind three of the other five. It produces 0.86, versus 0.94 for simple Markowitz and 0.99 for both 60/40 variants.

Our run was poor

We implemented the simple-Markowitz variant using EWMA alpha with a 252-day half-life scaled by 21, annualized 11-day covariance, the 7% volatility ceiling and the L1 trust region around 50/30/20. Trades execute at month-end MOC. Every figure in this section is ours. Our period runs from January 2020 to July 2024, roughly 4.5 years compared with the paper's twenty.

Our portfolio delivered 3.93% CAGR at 9.81% volatility and Sharpe 0.43. Over 2006 to 2026, the paper gives the same variant a 10.4% return, 9.3% volatility and Sharpe 0.91. Different windows separate the results, but ours is lower by 6.5 points of return and 0.48 of Sharpe. Our ratio also resembles raw return divided by volatility, 3.93/9.81 = 0.40, more closely than the paper's excess-of-fed-funds convention. Fed funds averaged around 2.5% during our period and exceeded 5%. Applying their convention would push our result well below 0.43. The observed gap is therefore a floor.

The risk machinery carried over. Reactive covariance. A hard ceiling. Realized volatility of 9.81% finished half a point above the paper's 9.3%. Our maximum drawdown was 16.38%, close to their 15.8% for simple Markowitz, although the respective windows are 2020 to 2024 and 2006 to 2026. The failure came from returns.

We verified two causes. Our 4.5-year sample begins with the February to March 2020 crash, when a trailing forecast with a 252-day half-life entered maximally long equities. It also includes the 2022 joint decline in stocks, bonds and gold. A momentum tilt had no defensive asset available there, while an 11-day covariance estimate could de-risk only after the move. Our SPY beta was 0.28, compared with the paper's 46.5% average SPY weight for simple Markowitz. The portfolio spent much of this stressed interval de-risked. Both episodes reduced returns while realized volatility remained near target.

The second cause is a fidelity gap that belongs to us. Our optimizer carries the paper's constraints without its equation (3) objective, the alpha-plus-cash-minus-turnover term. Those constraints fail to determine a unique weight vector. Our allocations can therefore drift from the return-maximizing corner of the feasible set, reducing return even as the volatility ceiling continues to bind.

Our costs are also higher. We charge the same 5bp spread, plus commissions of $0.0040 per share subject to a $1.00 order minimum, on roughly 280% annualized turnover. All reported figures are net of these costs. MOC orders fill at the auction print, giving zero slippage. The paper's cost sensitivity moves Sharpe from 1.08 to 1.03 when the spread quadruples to 20bp, suggesting a limited drag. We use monthly FEDFUNDS for cash in place of the paper's daily DFF, and our series is flagged non-vintage. The direction is ambiguous and the magnitude small at average cash weights of 13 to 15%.

We cannot fully account for the gap. The sample window is a large, verifiable contributor. With the objective specification unresolved, however, we cannot assign the remainder solely to regime. Our run says more about our implementation than it does about theirs.

Evidence that would change my view

I would want a covariance window chosen before the years used to judge it. The authors approximate this with annual walk-forward reselection over a grid that includes the covariance window, using deployment periods of 2011 to 2026 and 2016 to 2026. Retuning loses. The fixed specification beats walk-forward selection 1.00 to 0.93 under a trailing five-year rule, and 1.08 to 1.00 under a trailing ten-year rule.

The result supports the fixed choice's persistence under retuning. It also leaves the 11-day window fixed with knowledge of the entire 2006 to 2026 sample, while the only forward test of retuning favours that fixed specification. The paper's sample ends at its writing date, leaving no post-selection years. Its deflated Sharpe exceeds 0.99, based on 17 documented one-at-a-time neighbours whose trial Sharpe dispersion is 0.07. The authors concede that these neighbours are highly correlated, making the deflation likely understated.

A narrower conclusion than the abstract offers is still useful. Diluting a fixed-weight benchmark toward cash with a short volatility estimate roughly halves maximum drawdown. For 60/40, drawdown falls from 33.7% to 16.9%; for 50/30/20, it falls from 27.1% to 15.9%. Sharpe rises from 0.56 to 0.71 on 60/40 and from 0.70 to 0.82 on 50/30/20. Returns pay for the protection, dropping from 8.1% to 7.1% and from 9.1% to 7.8%. Turnover climbs from 3.4% and 4.0% to 85.0% and 79.0%. Cash earns the effective federal funds rate, which the paper concedes retail investors cannot obtain.

I would act on that result. Replacing momentum alpha with the 42-feature ridge regression adds 0.17 to the point estimate, taking Sharpe from 0.91 to 1.08. The return-level bootstrap places that difference in [-0.06, +0.40] with P = 0.068. The end-to-end bootstrap cannot separate it from zero. Against 60/40, the return-level bootstrap supports the authors' case at P = 0.011, while the end-to-end resample leaves it unresolved at P = 0.280. Turnover reaches 284.4%, versus 79.0% for the volatility overlay. Their own lot-level model, including wash sales and the 28%-capped collectibles rate on GLD, gives a top-bracket post-tax Sharpe of 0.64 against 0.53. Choosing the regression portfolio over the one-line cash dilution rule in equation (1) remains a live decision.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

Universe = [SPY, AGG, GLD] plus non-traded CASH

For each trading day:
  1. Compute adjusted-close simple returns for each ETF.
  2. Accrue CASH using the latest eligible FEDFUNDS observation:
       cash_return = (1 + FEDFUNDS / 100)^(1/252) - 1
  3. Apply the day's returns to existing holdings and update drifted weights.

On the last trading day of each month, after step 3:
  4. Estimate monthly expected returns as 21 times the EWMA of historical
     daily returns, using a 252-trading-day half-life.
  5. Estimate annualized covariance as 252 times the sample covariance of
     the latest 11 synchronized daily returns.
  6. If fewer than 11 complete synchronized returns exist, skip rebalance.
  7. Select risky weights w for SPY, AGG, and GLD using the intended
     21-day return-minus-cash-minus-turnover objective, subject to:
       w_i >= 0
       sum(w) <= 1
       sqrt(w' Sigma w) <= 0.07
       ||w - sum(w) * [0.5, 0.3, 0.2]||_1 <= sum(w)
  8. Set CASH weight = 1 - sum(w).
  9. Trade at that day's MOC print after charging applicable costs.
     New weights begin earning returns on the next trading day.

Track daily portfolio value, drifted weights, turnover, and gross/net returns.