AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Timing absorbs a third of SSO's shortfall

Bianchi and Goldberg study 501 days of levered ETF returns, with fees and NAV gaps left in the residual

2026-09-15 · 6 min read · US ETFs

Reviewing: A levered ETF anomaly explained · Stephen W. Bianchi and Lisa R. Goldberg · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

Conservative Timing-Covariance Volatility-Drag Allocation for SSO and UPRO

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (3)

  • Daily arithmetic return R_t = (P_{t+1} - P_t) / P_t. (invalidates: Exact reproduction of all reported cumulative, geometric, arithmetic, volatility, return-ratio and covariance results in Tables 1-5)
  • Daily return ratio lambda_ETF_t = R_ETF_t/R_t. (invalidates: Figure 4 clouds around the reported levels; Table 3 average ratios, ratio standard deviations and covariances; Table 5 winsorized ratios, ratio standard deviations and covariances)
  • Do not silently substitute dividend-inclusive total returns for the paper's price-return construction. (invalidates: Exact reproduction of Tables 1-5 and Figures 1-4)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 0.3%Jul 2024
Sharpe
0.05
Total Return
0.3%
Max Drawdown
-3.2%
CAGR
0.1%
Volatility
1.4%
Beta vs SPY
0.01
Trades
6

Begin with the return ratio

SSO delivered less than scaled index returns even though its average measured leverage exceeded 2.0. For each of the 501 trading days from January 3, 2022 to December 29, 2023, divide SSO's price return by the S&P 500's price return. The resulting daily return ratio averaged 2.033. UPRO averaged 3.106 against its stated target of 3.0. Realized returns went in the opposite direction. SSO's annualized arithmetic return was 1.655%, while twice the index's 1.917% comes to 3.833%. Applying the measured ratio of 2.033 gives 3.897%. UPRO earned 0.309%, versus 5.750% for a clean 3x and 5.952% using its average measured ratio of 3.106.

A product's average differs from the product of its averages by a sample covariance between the ratio and the index return. Bianchi and Goldberg report an annualized -2.242 percentage points for SSO and -5.643 for UPRO. Their name for this term is timing.

The paper asks why the ratio should contain information in the first place. Its answer starts with the measurement mismatch. A daily-levered fund targets a multiple of the index return calculated from net asset value, while the investor receives a multiple of the fund's price. The authors identify one possible source as the gap between net asset value, the market value of the underlying holdings, and the ETF price. They also state that the statistic "is not intended to recover portfolio or target return leverage, but rather to measure empirical return leverage from the perspective of investors." Calling the remaining term timing therefore gives a mechanism's name to what could have remained a residual.

Where the returns went

From January 3, 2022 to December 29, 2023, the S&P 500 price index gained +0.076% cumulatively and VOO gained +0.053%. SSO lost -11.088%. UPRO lost -28.238%. These are all CRSP price returns, with dividends excluded from both sides.

Volatility drag supplies the standard account. The paper measures it through a second-order Taylor approximation that links geometric return with arithmetic mean and variance. Start with the index's 1.917% mean and 19.40% volatility, scale each by 2x and 3x, and apply the formula. Estimated annualized geometric return becomes -3.625% for SSO and -10.580% for UPRO. The actual figures were -5.707% and -15.288%.

Our accounting uses 2x the index's realized geometric return, 0.076%, as the naive reference. The paper's own ideal 1x estimate is 0.035%. For SSO, the volatility term consumes 3.70 points, with another 2.08 points missing. UPRO loses 10.69 points to volatility and leaves 4.71 unexplained.

The abstract assigns roughly two-thirds of the return difference between the index and the levered ETFs to the familiar long-horizon volatility effect. The remaining amounts are 2.08 points for SSO and 4.71 for UPRO. Those are close to the paper's covariance estimates of -2.242 and -5.643, though they come from separate calculations. Table 3 contains the covariances in an arithmetic-return decomposition. The residual figures come from geometric returns. The tables answer different questions.

Arithmetic return divided by volatility was 0.099 for the index, 0.098 for VOO, 0.043 for SSO and 0.005 for UPRO. VOO retains the index's risk-adjusted return. The levered funds give it up.

The gap already appears in daily returns.

An identity without a mechanism

E(R_ETF) = E(lambda)E(R) + Cov(lambda, R) follows algebraically for any pair of return series once lambda is defined as their ratio. The covariance collects the distance between the fund's realized average return and the return implied by its average measured leverage. Fund expenses enter there, along with swap or futures financing. Distributions removed through the use of price returns also enter, as do deviations between NAV and price. The paper derives the identity while "neglecting dividends and market frictions," and its data appendix confirms that every series uses price returns. The reported -2.242 and -5.643 do not separate those components.

VOO exposes the problem. An unlevered S&P tracker has no leverage decision to mistime, yet its average measured return ratio is 1.028, with ratio volatility of 20.213 and covariance of -0.073. After the ratios are winsorized at the 95th percentile, VOO's covariance changes sign to +0.250 and its average ratio drops to 0.983. A mild upper-tail trim therefore moves a benchmark tracker from -0.073 to +0.250. The statistic is picking up estimation noise when the denominator is small. The authors identify the source directly: "The most extreme outliers occur when the index returns are close to zero."

SSO and UPRO are much less sensitive to that trim. Table 5 shows SSO's covariance moving from -2.242 to -2.215 and UPRO's from -5.643 to -5.100. Winsorization at the 95th percentile affects roughly 25 of the 501 days. Their shortfall against scaled index returns persisted through the sample instead of residing in the top 5% of return-ratio days. Persistence leaves the source unresolved between timing and cost.

A second geometric-return table adds two estimate rows. Its Empirical Estimate is -5.694% for SSO, against an actual -5.707%, and -15.226% for UPRO, against -15.288%. The paper offers this row as evidence that the decomposition works. Because the calculation uses the ETFs' realized mean and volatility, the fit is in-sample by construction.

We found no standard errors or confidence intervals for any covariance term. The evidence comes from a single 501-day window covering one index family.

We put the forecast into a rule

The paper ends with the forward-looking issue in its own words: "Within this framework, forecasting long-horizon levered ETF returns requires a view on the covariance term." It answers by identifying the four inputs required by the decomposition. We used those inputs to build a trading rule.

The following figures are ours. The paper remains descriptive and reports no strategy performance. Our run spans 2020-01-01 to 2024-07-01. At each month-end, we use only information available through the previous close. Over a strictly lagged 501-day window, matching the length of the paper's sample, we estimate the average return ratio, its covariance with VOO returns and the fund's volatility.

We calculate the same quantities again after capping ratios at the window's 95th percentile, then retain the worse result. A 500-replication stationary block bootstrap uses block length 20. From that exercise, we subtract 1.645 standard errors from the covariance and from the final geometric forecast. We also inflate estimated leveraged volatility by 1.25x.

A penalized forecast above VOO's triggers a position in SSO or UPRO, capped at one third of capital, with the balance held in cash. A 15% drawdown from the entry high or 45% annualized volatility during the preceding 20 days forces an exit to cash until the next month-end. Trades fill at the closing auction. Commissions are $0.0040 per share, subject to a $1 minimum.

Over 2020-01-01 to 2024-07-01, our rule returned 0.30%, with a Sharpe of 0.05, Sortino of 0.01 and Calmar of 0.02. Max drawdown was -3.20%, and annualized volatility was 1.37%.

Flat.

Two design choices dominate that result. With a 501-day lookback, the first tradable signal appears around the beginning of 2022. The live decision set is therefore roughly thirty monthly choices, almost all within the paper's 2022-2023 sample. The 45% volatility stop also falls below UPRO's realized annualized volatility of 58.01% during the paper's window. By construction, the overlay removes UPRO soon after most entries.

The one-third position cap, with the remaining capital assigned a zero assumed cash return, suppresses the result further. A portfolio invested by at most one third and making about thirty decisions has little scope to move. A 0.30% total return at 1.37% volatility is the outcome. These figures mainly reflect our design and carry no verdict on the paper. We did not model swap financing or the funds' expense ratios, which occupy the same unresolved space inside the paper's covariance term.

Stability remains untested

The decomposition provides clean ex post attribution, and its figures are internally consistent. Forward-looking use depends on a stable covariance term. The study examines one 501-day window, one index and three ETFs, supplies no error bars, and uses a statistic that changes sign for an unlevered tracker. It reports no evidence from other levered ETF families or disjoint two-year windows. It also leaves the funds' stated expense ratios and financing spreads inside the residual before calling that residual timing.

Until those tests are run, the covariance term remains whatever portion of return the volatility formula failed to capture.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For every trading day:
    Use only prices and returns completed by the previous close.

At the last trading day of each month:
    Require 501 synchronized lagged daily returns for VOO, SSO, and UPRO.
    For each leveraged ETF:
        Compute lambda_t = ETF_return_t / VOO_return_t when VOO_return_t != 0.
        Estimate raw mean lambda, covariance(lambda, VOO return), and ETF volatility.
        Separately cap lambda at its window-specific 95th percentile and recompute.
        Run 500 stationary-block bootstrap replications with block length 20 and seed 17.
        Subtract 1.645 bootstrap standard errors from raw and winsorized covariance.
        Multiply ETF volatility by 1.25.
        Convert raw and winsorized arithmetic forecasts to geometric forecasts.
        Set robust forecast to min(raw forecast, winsorized forecast).
        Subtract 1.645 bootstrap standard errors from the robust forecast.
    Compute VOO's volatility-drag-adjusted geometric forecast.
    Rank SSO and UPRO by penalized forecast.
    If the best leveraged forecast is strictly above VOO's forecast:
        Hold that ETF at min(33.3333%, weight allowed by max leverage); hold remainder in cash.
    Else if VOO's forecast is strictly above the 0 daily cash forecast:
        Hold VOO under the same cap; hold remainder in cash.
    Else:
        Hold cash.
    Execute changes at the current close using MOC pricing.

On each intervening trading day while a risky asset is held:
    From information through the previous close, calculate:
        drawdown = lagged close / highest lagged close since entry - 1
        volatility = sqrt(252) * population SD of the prior 20 completed returns
    If drawdown reaches -15% or annualized volatility exceeds 45%:
        Exit at the current close and remain in cash until the next monthly rebalance.

If an execution price is missing:
    Skip the order; never synthesize or forward-fill the price.