PEP-KO's $28,131 out-of-sample loss comes from one trade that stayed open until the data ran out. Graziano makes a stronger case for distrusting the profitable backtest than for abandoning the pair. Its best stretch depends heavily on one volatility episode, while the losing policy has no stop.
What Graziano trades in PEP-KO
The trade takes opposing positions in a cointegrated spread, aiming to collect when a divergence closes. Over 2013 to 2018, PepsiCo and Coca-Cola daily log returns have a correlation of 0.66; their annualised volatilities are 13.09% and 13.80%. If a linear combination of their log prices is stationary, a trader can short the rich leg and buy the cheap one, seeking convergence without taking market direction.
Graziano regresses log PEP on log KO using 2013 to 2018 data and estimates a hedge ratio of 1.662. With adjusted critical values, the Engle-Granger test rejects no-cointegration at p=0.033. The residual spread's AR(1) coefficient is 0.9785, implying a half-life of about 31.9 trading days.
The policy freezes those estimates. It expresses the spread as a z-score using its 2013-2018 mean and standard deviation, then shorts when z crosses above an entry level: one unit of PEP short against 1.662 units of KO long. Below minus that level, it goes long. A position closes when |z| falls below an exit level. Position size divides a 0.5% risk budget by the spread's standard deviation and the trailing 64-day volatility of daily z changes, with a 25% of equity cap. The optimiser deducts costs of 0.20% of traded notional on entry and exit of each leg.
Across 901 threshold pairs tested for maximum net Sharpe on 2018 to 2023, entry 2.15 and exit 0.50 win. The result is 9 round trips, a 0.67 annualised net Sharpe and $16,941 on $100,000.
Those settings stay fixed from 2023 to the present. One trade produces a -1.19 Sharpe and a $28,131 net loss. The abstract attributes the failure plainly to weakening mean reversion in the spread. The conclusion allows less certainty: the small trade count and limited test power prevent a definitive conclusion.
Graziano has evidence beyond the loss. The hedge ratio changes sign, from 1.6618 to -0.2376; a post-hoc Engle-Granger test gives p=0.12; and the adaptive variants lose too. The sign change matters. The dollar losses carry less weight when every out-of-sample run ends in forced liquidation under an exit rule with no stop, a point the paper never discusses.
What survives without COVID?
Sharpe drops from 0.67 to 0.28. Graziano removes the COVID window, located in February to April 2020, and reruns the pre- and post-COVID segments with separate volatility buffers. The optimiser again chooses (2.15, 0.50), but trades decline from 9 to 8 and net P&L from $16,941 to $5,759. The paper's description of the edge outside acute volatility regimes is apt: it "appears modest at best."
On the ex-COVID sample, the Pézier-White adjusted Sharpe is 0.281. The deflated Sharpe (DSR) is 0.423; it tests whether the best of 901 zero-skill policies could match the observed result. Graziano calls the statistical evidence of skill limited. I agree.
The walk-forward reaches a similar result by testing settings on the next period after each training window. Three folds each train on two years, producing an aggregate out-of-sample Sharpe of 0.64. Fold 1 tests on 2020, chooses the tighter (1.30, 0.65), trades 12 times and makes $13,794 at a 2.14 Sharpe. Folds 2 and 3 each trade once, losing $391 and $2,796. The 0.64 comes from the COVID year again.
Those test years, 2020 through 2022, fall within the 2018-2023 window used to select the headline thresholds. The walk-forward checks the procedure, yet shares data with the 0.67 result.
I give the sensitivity table less weight. Entry and exit thresholds remain the winners at costs from 0.05% to 0.50% and volatility windows of 6 days or more. With nine trades, the same few positions could dominate every cell and leave the argmax unchanged. J* still moves from 0.32 to 0.77 across the cost grid.
One position, held to the end
The frozen policy enters early in the out-of-sample window. The spread never returns inside 0.50, equity falls almost monotonically from a few weeks after entry, and the paper liquidates the position when the sample ends. This is thin evidence for the out-of-sample P&L claim, as Graziano acknowledges.
He says the -1.19 Sharpe is not a meaningful estimate of OOS performance and that one trade "does not allow a definitive attribution." Instead, he points to a hedge ratio re-estimated on the OOS window: -0.2376 against 1.6618. The post-hoc Engle-Granger test no longer rejects at p=0.12. A sign flip between two beverage peers is a real finding. The size of the dollar loss also reflects the exit rule.
As described, a position leaves through a z_exit crossing or end-of-sample liquidation. I found no stop-loss or time stop. With a 31.9-day half-life, one year amounts to about eight half-lives. After a few pass without reversion, the spread is already challenging the model.
The frozen z-score remains centred on the 2013-2018 mean and scaled by the 2013-2018 standard deviation. Drift in the spread level could therefore keep the 0.50 exit band out of reach. The adaptive runs make that explanation less persuasive. The $28,131 loss also depends on an end date identified only as the present.
Adaptive hedges, unchanged exit
Motivated by the sign flip, Graziano tries adaptive variants that re-standardise z on trailing 504-day data. A 504-day rolling OLS hedge ratio leaves the strategy flat for 18 months before one trade opens in mid-2024. Forced liquidation closes it at a $27,163 loss and a -1.05 Sharpe. The Kalman filter makes five trades, loses $32,752 and records a -0.88 Sharpe, worse in dollars than the static model. The paper says adaptation "does not improve OOS performance." It treats both variants as diagnostic because neither received walk-forward or DSR treatment.
Both variants re-estimate the intercept and hedge ratio and use a trailing 504-day mean and standard deviation for z. Yet each finishes with a position that never reaches the 0.50 exit. Stale standardisation cannot explain that result on its own. The absence of a stop remains.
The rolling beta moves from about 1 to -0.5, then partly recovers. The relationship moved. The P&L says less about the consequences because both variants retain the (2.15, 0.50) thresholds and the hold-until-exit rule. Rolling OLS loses on a position still open at sample end; the Kalman run's 5 trades likewise finish with forced liquidation.
A fixed time exit would change my view if the losses survived it. Rerun OOS with a forced exit at a fixed multiple of the 31.9-day half-life. Losses in both frozen and adaptive versions would give the break argument support beyond trades held to an arbitrary end date.
Our 0.11 Sharpe and its limits
We built a version from the paper's description and ran it over 2020-01-01 to 2024-07-01. These are our figures: 0.11 Sharpe, +0.18% total return, 0.37% annualised volatility, -1.37% max drawdown and 60 trades. Across those 60 trades, winners exceed losers by a profit factor of just 1.07. Graziano reports 0.67 in-sample for 2018-2023 and -1.19 out-of-sample for 2023-present. Our 0.11 lies between them, though our window matches neither period. Neither difference is a like-for-like comparison.
Our run contains the COVID months responsible for much of the paper's in-sample P&L. It covers only 18 months of the paper's OOS period, likely missing much of the loss Graziano describes. His rolling-OLS variant stays flat for its first 18 months and opens its sole trade in mid-2024, near the end of our window.
We used 2013-2017 for formation and 2018-2022 for selection. Both windows end a year ahead of the paper's 2013-2018 and 2018-2023 windows, likely changing fitted parameters and possibly the selected thresholds. We charged $0.004 a share with a $1 minimum. At these prices, that is far cheaper than 0.20% of notional and flatters our result. Fills were market-on-close, although we had intended to fill at the next session's open.
The 0.37% volatility deserves particular attention. Graziano's OOS P&L moves by $28,131 on $100,000, while our realised volatility suggests effective exposure at a small fraction of his. Our per-instrument caps and proportional leg reductions are likely part of the reason. We cannot fully account for the scale gap from what we can see.
Per-fill counting might explain the 60 trades. A different signal path might too; we could not determine which. This was one automated pass, and its result says more about our implementation than about Graziano's work.
Graziano is candid about 0.67 falling to 0.28. I would wait for losses under a time stop before writing PEP-KO's obituary.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.