Begin with the return ratio
SSO delivered less than scaled index returns even though its average measured leverage exceeded 2.0. For each of the 501 trading days from January 3, 2022 to December 29, 2023, divide SSO's price return by the S&P 500's price return. The resulting daily return ratio averaged 2.033. UPRO averaged 3.106 against its stated target of 3.0. Realized returns went in the opposite direction. SSO's annualized arithmetic return was 1.655%, while twice the index's 1.917% comes to 3.833%. Applying the measured ratio of 2.033 gives 3.897%. UPRO earned 0.309%, versus 5.750% for a clean 3x and 5.952% using its average measured ratio of 3.106.
A product's average differs from the product of its averages by a sample covariance between the ratio and the index return. Bianchi and Goldberg report an annualized -2.242 percentage points for SSO and -5.643 for UPRO. Their name for this term is timing.
The paper asks why the ratio should contain information in the first place. Its answer starts with the measurement mismatch. A daily-levered fund targets a multiple of the index return calculated from net asset value, while the investor receives a multiple of the fund's price. The authors identify one possible source as the gap between net asset value, the market value of the underlying holdings, and the ETF price. They also state that the statistic "is not intended to recover portfolio or target return leverage, but rather to measure empirical return leverage from the perspective of investors." Calling the remaining term timing therefore gives a mechanism's name to what could have remained a residual.
Where the returns went
From January 3, 2022 to December 29, 2023, the S&P 500 price index gained +0.076% cumulatively and VOO gained +0.053%. SSO lost -11.088%. UPRO lost -28.238%. These are all CRSP price returns, with dividends excluded from both sides.
Volatility drag supplies the standard account. The paper measures it through a second-order Taylor approximation that links geometric return with arithmetic mean and variance. Start with the index's 1.917% mean and 19.40% volatility, scale each by 2x and 3x, and apply the formula. Estimated annualized geometric return becomes -3.625% for SSO and -10.580% for UPRO. The actual figures were -5.707% and -15.288%.
Our accounting uses 2x the index's realized geometric return, 0.076%, as the naive reference. The paper's own ideal 1x estimate is 0.035%. For SSO, the volatility term consumes 3.70 points, with another 2.08 points missing. UPRO loses 10.69 points to volatility and leaves 4.71 unexplained.
The abstract assigns roughly two-thirds of the return difference between the index and the levered ETFs to the familiar long-horizon volatility effect. The remaining amounts are 2.08 points for SSO and 4.71 for UPRO. Those are close to the paper's covariance estimates of -2.242 and -5.643, though they come from separate calculations. Table 3 contains the covariances in an arithmetic-return decomposition. The residual figures come from geometric returns. The tables answer different questions.
Arithmetic return divided by volatility was 0.099 for the index, 0.098 for VOO, 0.043 for SSO and 0.005 for UPRO. VOO retains the index's risk-adjusted return. The levered funds give it up.
The gap already appears in daily returns.
An identity without a mechanism
E(R_ETF) = E(lambda)E(R) + Cov(lambda, R) follows algebraically for any pair of return series once lambda is defined as their ratio. The covariance collects the distance between the fund's realized average return and the return implied by its average measured leverage. Fund expenses enter there, along with swap or futures financing. Distributions removed through the use of price returns also enter, as do deviations between NAV and price. The paper derives the identity while "neglecting dividends and market frictions," and its data appendix confirms that every series uses price returns. The reported -2.242 and -5.643 do not separate those components.
VOO exposes the problem. An unlevered S&P tracker has no leverage decision to mistime, yet its average measured return ratio is 1.028, with ratio volatility of 20.213 and covariance of -0.073. After the ratios are winsorized at the 95th percentile, VOO's covariance changes sign to +0.250 and its average ratio drops to 0.983. A mild upper-tail trim therefore moves a benchmark tracker from -0.073 to +0.250. The statistic is picking up estimation noise when the denominator is small. The authors identify the source directly: "The most extreme outliers occur when the index returns are close to zero."
SSO and UPRO are much less sensitive to that trim. Table 5 shows SSO's covariance moving from -2.242 to -2.215 and UPRO's from -5.643 to -5.100. Winsorization at the 95th percentile affects roughly 25 of the 501 days. Their shortfall against scaled index returns persisted through the sample instead of residing in the top 5% of return-ratio days. Persistence leaves the source unresolved between timing and cost.
A second geometric-return table adds two estimate rows. Its Empirical Estimate is -5.694% for SSO, against an actual -5.707%, and -15.226% for UPRO, against -15.288%. The paper offers this row as evidence that the decomposition works. Because the calculation uses the ETFs' realized mean and volatility, the fit is in-sample by construction.
We found no standard errors or confidence intervals for any covariance term. The evidence comes from a single 501-day window covering one index family.
We put the forecast into a rule
The paper ends with the forward-looking issue in its own words: "Within this framework, forecasting long-horizon levered ETF returns requires a view on the covariance term." It answers by identifying the four inputs required by the decomposition. We used those inputs to build a trading rule.
The following figures are ours. The paper remains descriptive and reports no strategy performance. Our run spans 2020-01-01 to 2024-07-01. At each month-end, we use only information available through the previous close. Over a strictly lagged 501-day window, matching the length of the paper's sample, we estimate the average return ratio, its covariance with VOO returns and the fund's volatility.
We calculate the same quantities again after capping ratios at the window's 95th percentile, then retain the worse result. A 500-replication stationary block bootstrap uses block length 20. From that exercise, we subtract 1.645 standard errors from the covariance and from the final geometric forecast. We also inflate estimated leveraged volatility by 1.25x.
A penalized forecast above VOO's triggers a position in SSO or UPRO, capped at one third of capital, with the balance held in cash. A 15% drawdown from the entry high or 45% annualized volatility during the preceding 20 days forces an exit to cash until the next month-end. Trades fill at the closing auction. Commissions are $0.0040 per share, subject to a $1 minimum.
Over 2020-01-01 to 2024-07-01, our rule returned 0.30%, with a Sharpe of 0.05, Sortino of 0.01 and Calmar of 0.02. Max drawdown was -3.20%, and annualized volatility was 1.37%.
Flat.
Two design choices dominate that result. With a 501-day lookback, the first tradable signal appears around the beginning of 2022. The live decision set is therefore roughly thirty monthly choices, almost all within the paper's 2022-2023 sample. The 45% volatility stop also falls below UPRO's realized annualized volatility of 58.01% during the paper's window. By construction, the overlay removes UPRO soon after most entries.
The one-third position cap, with the remaining capital assigned a zero assumed cash return, suppresses the result further. A portfolio invested by at most one third and making about thirty decisions has little scope to move. A 0.30% total return at 1.37% volatility is the outcome. These figures mainly reflect our design and carry no verdict on the paper. We did not model swap financing or the funds' expense ratios, which occupy the same unresolved space inside the paper's covariance term.
Stability remains untested
The decomposition provides clean ex post attribution, and its figures are internally consistent. Forward-looking use depends on a stable covariance term. The study examines one 501-day window, one index and three ETFs, supplies no error bars, and uses a statistic that changes sign for an unlevered tracker. It reports no evidence from other levered ETF families or disjoint two-year windows. It also leaves the funds' stated expense ratios and financing spreads inside the residual before calling that residual timing.
Until those tests are run, the covariance term remains whatever portion of return the volatility formula failed to capture.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.