Lag the return forecast by a full month and Sharpe slips from 1.08 to 1.00. Lag both the return forecast and covariance estimate by that month, as the paper's Table 12 does, and Sharpe drops to 0.73. Those appendix D results contain the paper's real economic result. The abstract leads with the volatility overlay, which the authors say improves every risk-adjusted and drawdown metric. Appendix D assigns more importance to covariance than alpha. The choice of the regression portfolio as the headline book therefore needs explaining.
Three ETFs, cash and a 7% cap
The investable set contains SPY, AGG, GLD and cash accruing at the effective federal funds rate. The portfolios are long only, use no leverage or derivatives, and rebalance monthly from public data. Open-source code reproduces every result. The test runs from January 2006 to January 2026, twenty years, and charges a flat 5bp bid-ask spread on every trade. The optimizer anticipates that charge in its objective. The authors frame the contribution as empirical rather than methodological, accompanied by a reproducible implementation that a self-directed investor could operate.
Six portfolios make up the comparison. Two are fixed-weight benchmarks rebalanced annually: conventional 60/40 and a 50/30/20 stock/bond/gold allocation motivated by rising stock-bond correlation. Another two apply volatility control to those benchmarks. Every month, the allocation moves toward cash until estimated ex-ante volatility reaches 7% annualized. That estimate comes from the benchmark's own daily returns over an 11-trading-day trailing window.
The remaining two portfolios optimize weights. Each maximizes the same monthly objective: forecast return, plus the cash return earned by the unallocated share, minus an L1 turnover cost. Two constraints do the heavy lifting. Estimated portfolio volatility has a hard 7% ceiling, while an L1 trust region holds relative weights within distance 1 of (0.5, 0.3, 0.2). The portfolio is long-only, and weights sum to at most one. Covariance comes from the annualized 11-day trailing sample matrix.
Only the return forecast changes between them. Simple Markowitz takes an EWMA of past daily returns with a 252-day half-life, then scales it by 21. The full Markowitz portfolio predicts 100-day forward returns through a ridge regression using 42 features, lambda 10, at most 512 training observations and 252-day half-life weights. Inputs include trailing returns and volumes, Treasury yields, breakevens, smoothed VIX, the dollar index and eight Fama-French factors. Its final training target is dated t minus 100, preserving causality.
After the 5bp spread, Markowitz returns 11.6% with 9.0% volatility, Sharpe 1.08, an 18.1% maximum drawdown and 284.4% turnover. Simple Markowitz records 10.4%, 9.3% and Sharpe 0.91. The volatility-controlled 50/30/20 portfolio produces 7.8%, 7.3% and 0.82. Conventional 60/40 returns 8.1% at 11.3% volatility, with Sharpe 0.56 and a 33.7% drawdown. Both optimization portfolios stayed positive through 2008, returning +4.0% and +2.6%.
Where does the alpha come from?
The authors rule out the simplest account themselves. A static portfolio holding Markowitz's average weights, 41.5% SPY, 20.1% AGG, 23.3% GLD and 15.1% cash, returns 8.5% at 9.1% volatility when rebalanced annually. Its Sharpe is 0.73. Risk remains similar while risk-adjusted return falls to two thirds of the dynamic portfolio's figure. Variation supplies the advantage. In the authors' words, the benefit comes from "varying the cash holdings (as well as the mix of stocks, bonds, and gold) strategically."
Which variation matters?
The lag test separates the components. Delaying alpha by 21 trading days preserves 1.00 of the original 1.08. Delaying alpha and covariance together by 21 days leaves 0.73, below the 0.82 from the basic volatility-controlled 50/30/20 portfolio. A five-day delay to both reduces Sharpe to 0.81. The authors reach the same reading: a reactive covariance forecast "appears to be substantially more important than a reactive return forecast."
Appendix F makes the point harder to avoid. Seven standard risk-based rules use the same three assets, costs and constraints. Their volatility-controlled Sharpes range from 0.60 to 0.85. Each rule reaches its highest Sharpe with the 11-day covariance window selected for the authors' portfolio, and each deteriorates at 21, 63, 126 and 252 days. Global minimum variance declines from 0.60 to 0.37 with the 252-day window. The authors answer this evidence in the same appendix: all seven rules share the 11-day covariance estimate, yet none exceeds 0.85. They conclude that "sizing positions by risk is evidently not enough on its own; the lift comes from combining a return forecast with a hard limit on estimated risk."
Estimating a three by three covariance matrix from eleven daily observations is aggressive. The authors acknowledge this, describing the window as "short by the standards of volatility estimation" and the result as "a noisy estimate". Its appeal is speed, and they selected it "after a modest search over various values". The rationale makes sense. During 2008 and March 2020, a useful risk estimate needed to react within days. Yet the full sample was already visible when the window was chosen. The paper tabulates a covariance sweep from 11 to 252 days for the seven comparison rules, all peaking at 11 days. Appendix E gives only a qualitative account for the authors' own portfolio, observing that the covariance estimate window "does affect the performance, though smoothly."
The alpha model still deserves credit. Black-Litterman reaches just 0.77 with volatility control when given exactly the same regression forecast. Across 90 specifications swept in-sample, its best Sharpe is 0.95, short of 1.08. Those winning specifications set view confidence near zero and effectively disable the equilibrium prior. Here, the shrinkage Black-Litterman was designed to supply reduced performance.
The end-to-end bootstrap removes the margin
The paper reports two bootstraps.
First, it resamples portfolio returns using 21-day mean blocks, 10,000 replications and paired observations across portfolios. Markowitz beats 60/40 by +0.52 Sharpe, with interval [+0.07, +0.99] and P(difference at most zero) = 0.011. The result clears 5%. Against 50/30/20, the difference is +0.38 and the interval is [-0.01, +0.79]. Against simple Markowitz, it is +0.17 with [-0.06, +0.40] and P = 0.068. The headline portfolio therefore beats the classic benchmark significantly, while no other comparison reaches significance. The authors say this directly. They also explain that simple Markowitz and the volatility-controlled portfolios are their own constructions, so those tests help choose among their portfolios rather than establish superiority over the status quo. That framing is fair.
The second bootstrap carries more weight. It uses a mean block length of 300 days and 100 replications. Raw inputs are resampled against a common index, after which the features are rebuilt, the ridge is refitted and all six portfolios are rerun. The result is blunt: "Every interval for the Markowitz Sharpe advantage contains zero, and the estimated probability of a nonpositive difference ranges from 0.28 to 0.51." The mean difference against 50/30/20 VC is -0.01. Against 60/40, it is +0.17 with interval [-0.23, +0.64] and P = 0.280. The authors describe this as a conservative check based on 100 replications, rather than a precise tail calculation. They are right, and they publish the result anyway, which is uncommon.
The sample itself raises another concern. Its strongest five-year block includes 2008 and posts Sharpe 1.34. Removing 2008 and 2009 almost preserves the ranking, at 1.11, 0.95, 0.93, 0.83, 0.87 and 0.74, though Markowitz's lead over 50/30/20 VC narrows from 0.26 to 0.18. During the 2011 to 2016 subperiod, the flagship portfolio ranks behind three of the other five. It produces 0.86, versus 0.94 for simple Markowitz and 0.99 for both 60/40 variants.
Our run was poor
We implemented the simple-Markowitz variant using EWMA alpha with a 252-day half-life scaled by 21, annualized 11-day covariance, the 7% volatility ceiling and the L1 trust region around 50/30/20. Trades execute at month-end MOC. Every figure in this section is ours. Our period runs from January 2020 to July 2024, roughly 4.5 years compared with the paper's twenty.
Our portfolio delivered 3.93% CAGR at 9.81% volatility and Sharpe 0.43. Over 2006 to 2026, the paper gives the same variant a 10.4% return, 9.3% volatility and Sharpe 0.91. Different windows separate the results, but ours is lower by 6.5 points of return and 0.48 of Sharpe. Our ratio also resembles raw return divided by volatility, 3.93/9.81 = 0.40, more closely than the paper's excess-of-fed-funds convention. Fed funds averaged around 2.5% during our period and exceeded 5%. Applying their convention would push our result well below 0.43. The observed gap is therefore a floor.
The risk machinery carried over. Reactive covariance. A hard ceiling. Realized volatility of 9.81% finished half a point above the paper's 9.3%. Our maximum drawdown was 16.38%, close to their 15.8% for simple Markowitz, although the respective windows are 2020 to 2024 and 2006 to 2026. The failure came from returns.
We verified two causes. Our 4.5-year sample begins with the February to March 2020 crash, when a trailing forecast with a 252-day half-life entered maximally long equities. It also includes the 2022 joint decline in stocks, bonds and gold. A momentum tilt had no defensive asset available there, while an 11-day covariance estimate could de-risk only after the move. Our SPY beta was 0.28, compared with the paper's 46.5% average SPY weight for simple Markowitz. The portfolio spent much of this stressed interval de-risked. Both episodes reduced returns while realized volatility remained near target.
The second cause is a fidelity gap that belongs to us. Our optimizer carries the paper's constraints without its equation (3) objective, the alpha-plus-cash-minus-turnover term. Those constraints fail to determine a unique weight vector. Our allocations can therefore drift from the return-maximizing corner of the feasible set, reducing return even as the volatility ceiling continues to bind.
Our costs are also higher. We charge the same 5bp spread, plus commissions of $0.0040 per share subject to a $1.00 order minimum, on roughly 280% annualized turnover. All reported figures are net of these costs. MOC orders fill at the auction print, giving zero slippage. The paper's cost sensitivity moves Sharpe from 1.08 to 1.03 when the spread quadruples to 20bp, suggesting a limited drag. We use monthly FEDFUNDS for cash in place of the paper's daily DFF, and our series is flagged non-vintage. The direction is ambiguous and the magnitude small at average cash weights of 13 to 15%.
We cannot fully account for the gap. The sample window is a large, verifiable contributor. With the objective specification unresolved, however, we cannot assign the remainder solely to regime. Our run says more about our implementation than it does about theirs.
Evidence that would change my view
I would want a covariance window chosen before the years used to judge it. The authors approximate this with annual walk-forward reselection over a grid that includes the covariance window, using deployment periods of 2011 to 2026 and 2016 to 2026. Retuning loses. The fixed specification beats walk-forward selection 1.00 to 0.93 under a trailing five-year rule, and 1.08 to 1.00 under a trailing ten-year rule.
The result supports the fixed choice's persistence under retuning. It also leaves the 11-day window fixed with knowledge of the entire 2006 to 2026 sample, while the only forward test of retuning favours that fixed specification. The paper's sample ends at its writing date, leaving no post-selection years. Its deflated Sharpe exceeds 0.99, based on 17 documented one-at-a-time neighbours whose trial Sharpe dispersion is 0.07. The authors concede that these neighbours are highly correlated, making the deflation likely understated.
A narrower conclusion than the abstract offers is still useful. Diluting a fixed-weight benchmark toward cash with a short volatility estimate roughly halves maximum drawdown. For 60/40, drawdown falls from 33.7% to 16.9%; for 50/30/20, it falls from 27.1% to 15.9%. Sharpe rises from 0.56 to 0.71 on 60/40 and from 0.70 to 0.82 on 50/30/20. Returns pay for the protection, dropping from 8.1% to 7.1% and from 9.1% to 7.8%. Turnover climbs from 3.4% and 4.0% to 85.0% and 79.0%. Cash earns the effective federal funds rate, which the paper concedes retail investors cannot obtain.
I would act on that result. Replacing momentum alpha with the 42-feature ridge regression adds 0.17 to the point estimate, taking Sharpe from 0.91 to 1.08. The return-level bootstrap places that difference in [-0.06, +0.40] with P = 0.068. The end-to-end bootstrap cannot separate it from zero. Against 60/40, the return-level bootstrap supports the authors' case at P = 0.011, while the end-to-end resample leaves it unresolved at P = 0.280. Turnover reaches 284.4%, versus 79.0% for the volatility overlay. Their own lot-level model, including wash sales and the 28%-capped collectibles rate on GLD, gives a top-bracket post-tax Sharpe of 0.64 against 0.53. Choosing the regression portfolio over the one-line cash dilution rule in equation (1) remains a live decision.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.