AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Four option factors survive 160 controls, while spreads lose money

Walter, Zimmer and Ulrich find in-sample tail pricing; the Sharpe gain from 3.35 to 3.74 is insignificant.

2026-09-28 · 8 min read · Option-implied equity factors · Optionable U.S. stocks

Reviewing: Taming the Option Factor Zoo: A High-Dimensional Analysis · Alexander Walter, Lukas Zimmer and Maxim Ulrich · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Thirty-Day Option Characteristics: Subsequent Stock-Return Decile Spreads

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper (3)

  • The 2004–2023 sample cannot be replicated: available options history begins around 2020.
  • Point-in-time S&P 500 membership and the paper's Jensen et al. characteristic dataset are unavailable. An annual top-capitalization screen and locally constructed equity factors can be tested, but this does not reproduce the paper's exact universe or 160-factor benchmark.
  • The paper's option-implied physical-moment variants use realized high-frequency moments unavailable at sub-minute frequency; test EOD risk-neutral characteristics instead.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (9)

  • Section 2.2 physical-moment scaling: P_τ = (P/Q_1d) × Q_τ.: Retained verbatim as a paper reference but do not compute physical-moment signals. (invalidates: Any replication claim for the paper's physical-moment factors.)
  • Eq. (1) SDF cross-sectional model: r̄ = γ_0 + Ĉ_g λ_g + Ĉ_h λ_h + ε.: Compare subsequent portfolio returns and attainable-control regression alphas; do not estimate λ_g. (invalidates: Paper Table 6 conditional SDF loadings and t-statistics.)
  • Eq. (2) inference residual: z_t = g_t − Π_{I_3} g_t.: Do not run the paper's third-stage inference LASSO. (invalidates: Paper Table 6 conditional SDF loading inference.)
  • Eq. (3) Stewart–Love redundancy index: RI_{Y|X} = Σ_{i=1}^{k} L²_{Y,i} × ρ²_i.: Do not estimate the paper's 20-versus-160-factor canonical correlations. (invalidates: Paper Table 7 shared-variance values and noise comparisons.)

5 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total -21.5%Jul 2024
Sharpe
-0.19
Total Return
-21.5%
Max Drawdown
-44.4%
CAGR
-5.2%
Volatility
25.5%
Beta vs SPY
0.06
Trades
10,278

What the paper reports for its own strategy

  • bakshiKurt long-short factor (value-weighted decile, S&P 500 constituents, 30-day maturity): -11.23% annualized mean, vol 12.84%, Sharpe -0.87, MDD -91.07%, skew -2.29, t = -3.25 (Newey–West, 6 lags); Feb 2004–Jul 2023, gross of transaction costs
  • ivs_convexity long-short factor: -14.78% annualized mean, vol 17.94%, Sharpe -0.82, MDD -96.94%, skew -3.59, t = -2.57; Feb 2004–Jul 2023, gross of transaction costs
  • ivs_smirk long-short factor: -11.17% annualized mean, Sharpe -0.59, MDD -93.05%, t = -1.96; Feb 2004–Jul 2023, gross
  • BTX_EQV_relRightTail_10 long-short factor: 5.04% annualized, Sharpe 0.28, t = 1.23; Feb 2004–Jul 2023, gross
  • RND_mu long-short factor: 8.26% annualized, Sharpe 0.27, t = 1.06; Feb 2004–Jul 2023, gross
  • FF5+momentum annualized alphas (Newey–West t): bakshiKurt -10.53% (t = -3.14), ivs_convexity -16.71% (t = -2.79), ivs_smirk -12.64% (t = -2.47), SVIX2_put_annual -9.07% (t = -2.08); Feb 2004–Jul 2023, gross

A priced option exposure can make a lousy long-short trade. In Walter, Zimmer and Ulrich's paper, risk-neutral kurtosis and implied-volatility convexity survive equity-factor controls, yet their decile portfolios lose 11.2% and 14.8% a year, gross, across 234 months. The distinction between the loading and the payoff governs how useful these results are to a desk.

What enters the test?

An option surface gives a forward-looking view of a stock's return distribution, including its tails. Earlier work finds that option-implied measures predict cross-sectional returns. Neuhierl et al. showed that option characteristics predict returns beyond firm characteristics. Walter, Zimmer and Ulrich set a narrower hurdle: after controlling for 160 equity factors, does a portfolio sorted on an option characteristic enter the stochastic discount factor?

The authors start with CBOE data on 7,126 optionable U.S. stocks. They fit implied-volatility surfaces using the kernel estimator of Ulrich et al. (2023), extract risk-neutral densities, and calculate 137 characteristics at a 30-day constant maturity. The set includes Bakshi-Kapadia-Madan moments, Bollerslev-Todorov-Xu jump and tail decompositions (the BTX_ family), IV smirk and convexity, and SVIX. P_implQ variants use a one-day realized-to-implied ratio to scale risk-neutral moments into proxies for physical moments. Each characteristic produces a value-weighted top-minus-bottom decile portfolio among S&P 500 constituents, with about 42 stocks per leg and monthly rebalancing. The comparison set comprises 152 similarly built factors from Jensen-Kelly-Pedersen characteristics and 8 French factors. The sample runs for 234 months, from February 2004 to July 2023.

For pricing, the authors use the Feng-Giglio-Xiu double-selection LASSO. One selection finds equity factors that help price the test assets; the other finds equity factors with covariances resembling the candidate option factor's. OLS on their union estimates λs, the candidate's SDF loading scaled to a unit beta, in basis points per month. Among 20 screened option factors, Eight have loadings significant at 5%. Four pass the Bonferroni threshold of |t| > 3.02:

The screen sees the same 234 months

Many of the 137 characteristics carry almost the same information. A VIF above 10 applies to 91% of them; 62% exceed 100. To reduce that overlap, the authors fit a gradient-boosted regressor to the equal-weighted average test-asset return in month t, using the 137 option factor returns from that same month as features. SHAP attributions assign credit for the fit. The top 20 account for 89.3% of total attribution.

The screen uses a contemporaneous fit over the full sample. The authors explicitly condition their tests on that choice. They also acknowledge that selecting 20 from 137 leaves a multiple-testing burden beyond Bonferroni and that their criterion favours factors with large market exposure. Their frequency results show why this matters: bakshiKurt enters in 3 of 70 window-maturity combinations, RND_kurt_eps in 3, and BTX_MRJI_20_annual in 1. P_implQ_var appears in 32 and BTX_RJV_k in 28. The authors call the selections "representatives of groups of related measures... rather than as uniquely identified signals." bakshiKurt, despite clearing Bonferroni, is chosen in only 3 of 70 runs. The other two rarely selected factors clear only the 5% threshold.

How can a losing spread have a positive loading?

Take ivs_convexity. Its λs is +420 bp a month, while its average return is -123.2 bp a month. Gross FF5-plus-momentum alpha is -16.71% a year (t = -2.79). For bakshiKurt, λs is +78.51 bp and the average return is -93.6 bp a month; alpha is -10.53% (t = -3.14). In 5 of the 8 significant cases, λs and the average return have opposite signs.

The test portfolios include 3x2 sorts on each option characteristic and size. Across those portfolios, λs prices covariance with the candidate factor after holding the selected equity controls fixed. A raw decile spread carries its other exposures as well. The authors point out a limit to that explanation: bakshiKurt and ivs_convexity have small six-factor loadings, so those six exposures do not account for the sign gap.

A positive λs therefore describes a premium for high-convexity covariance in the test assets. It does not give the decile spread a positive payoff. We did not find a portfolio in the paper that isolates the control-neutral exposure captured by λs. The available long-short has a -96.94% maximum drawdown and skew of -3.59.

An insignificant Sharpe gain

With the 20 option factors added, the in-sample maximum Sharpe rises from 3.35 to 3.74. The option factors alone reach 1.29. The joint GRS statistic for their 20 alphas is 1.25, with p = 0.25. The authors identify the problems: estimating a maximum Sharpe with 180 factors over 234 months biases it upward, and the GRS test has few residual degrees of freedom. They argue that a joint test of 20 can dilute a few relevant factors.

That argument has force. We did not find a joint test confined to the four Bonferroni survivors. Across option maturities, the Sharpe increase runs from 0.16 to 0.39; the GRS rejects only at 60 days (p = 0.0427).

The redundancy results say more about what these portfolios contain. The 160 equity factors account for 91.0% of option factor variance, versus roughly 69% for pure noise with the same dimensions. In the other direction, the 20 option factors account for 35.4% of equity factor variance, versus roughly 9% for noise. The first comparison is roughly 1.3 times noise (91.0% against 69%); the second is roughly 3.9 times noise (35.4% against 9%). A small set of option sorts thus reproduces a third of the equity zoo. Heavy RMW and momentum loadings fit that result: bakshi_X loads at -1.53 and -0.51, and SVIX2_put_annual at -1.47 and -0.52.

Jump tails across maturities, kurtosis at 30 days

The authors rerun the screen and pricing tests from 1 to 60 days. BTX_LJV_7_5 is selected at five maturities and has a positive loading each time. Its t-statistics include 3.09 at 3 days, 2.62 at 5 days, and 3.59 at 30 days; at 1 day it barely clears the conventional threshold, at 1.97. BTX_EQV_relTail_k is positive and significant at all three maturities where the screen selects it. Kurtosis shows up only at 30 days. ivs_convexity switches to a negative loading at 10 days (-163 bp, t = -1.95). P_implQ_var goes from +91 (t = 3.46) at 1 day to -261 (t = -2.40) at 30 days.

The treatment of distant strikes bears directly on these findings. The risk-neutral densities use IVs extrapolated linearly beyond observed strikes, and all four Bonferroni survivors measure tails or curvature. Those are the measures most exposed to the extrapolated wing. We did not find a sensitivity check for the extrapolation rule.

Our four-signal book, 2020 to 2024

We could not reproduce the paper's setup. Our options history starts around 2020, leaving its 2004-2023 sample beyond reach. We lack point-in-time S&P 500 membership and the Jensen-Kelly-Pedersen dataset, so we substituted an annual top-1,000 non-ADR capitalization screen. P_implQ variants require realized high-frequency moments we do not have; we used end-of-day risk-neutral characteristics instead. We ran none of the paper's pricing tests.

Our book combined the four Bonferroni survivors as value-weighted top-minus-bottom deciles on 30-day option characteristics, formed monthly. We charged $0.004 a share in commissions. The signal definitions were not independently checked against the source equations, including the time-varying jump cutoff. Fills occurred at the close, although we intended next-session open fills. This run measures raw decile spreads. It cannot test the paper's conditional SDF-loading claim.

From January 2020 to July 2024, our book returned a -0.19 Sharpe and a -5.24% CAGR, net of commissions. Annualized volatility was 25.51%, and maximum drawdown was -44.44%. For comparison, the paper reports 12.84% volatility for bakshiKurt and 17.94% for ivs_convexity, with drawdowns of -91.07% and -96.94%, over 19.5 years on S&P 500 names. Its gross single-factor Sharpes on S&P 500 names over 2004-2023 are -0.87 for bakshiKurt and -0.82 for ivs_convexity; BTX_LJV_7_5 and BTX_EQV_relTail_k come in at -0.18 and -0.13. Our result shares the negative sign and is far weaker than the two strongly negative Sharpes. Different methods, universes and windows make these different measurements.

We cannot fully explain the gap.

Combining two strongly negative signals with two near-zero ones likely moves the book's Sharpe toward zero, though we have no per-signal returns with which to size the effect. Our 4.6-year window overlaps the paper's by only 3.5 years and probably misses episodes behind its -2.29 and -3.59 skews. Unvalidated, noisier signal construction could weaken the spread. The broader 1,000-stock universe could shift its magnitude either way; close-versus-open timing could too. Costs exclude short borrow and financing, both relevant when shorting four bottom deciles of tail-heavy names. This was one quick automated pass, showing the negative sign and nothing more, rather than a verdict on the authors' work.

Which signals merit desk time?

Four of the 20 factors clear Bonferroni. The paper finds a few in-sample priced dimensions concentrated in jump tails, or, in the authors' words, "a few pricing-relevant dimensions rather than a new factor zoo." BTX_LJV_7_5 and BTX_EQV_relTail_k deserve the closer look: their loadings remain positive and significant wherever the screen selects them, five maturities for LJV_7_5 and three for relTail_k. Their raw decile spreads remain close to zero and insignificant. LJV_7_5 earns -4.42% a year (t = -0.73, or -36.8 bp a month); relTail_k earns -2.35% (t = -0.47, or -19.6 bp a month). The priced exposure and the spread payoff are separate quantities.

Kurtosis appears only at 30 days. Convexity is positive and significant at 1, 2 and 30 days, then negative at 10 days. Their spreads lose 11.2% and 14.8% a year gross. The authors themselves propose the next test: screen on one part of the sample and test on the remainder. A positive left-tail jump-variation loading above t = 3 in that hold-out would move me.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each year, identify the top 1,000 eligible non-ADR stocks by that year's screening-table capitalization.
At each month-end, independently validate the source definitions and numerical implementation of all four signals before labeling any result with them.
For each valid stock-month, use observed option chains and a contemporaneous forward to estimate a 30-day risk-neutral distribution; omit chains that cannot support the measure.
For each signal separately, rank eligible stocks into deciles. Value-weight the top decile long and bottom decile short, sizing each leg to one unit of initial capital.
The specification calls for entries, exits, and monthly share changes at the next session's observed adjusted stock open; skip fills without an observed execution bar.
Report subsequent leg and spread returns, coverage, turnover, and uncertainty separately by signal.