AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Deep kernel hedging gains survive a ReLU swap

Dupret, Hainaut and Motte beat deep hedging in most tests, but their ablation narrows the kernel claim.

2026-09-29 · 8 min read · Derivatives Hedging · US listed equity options and underlying US equities

Reviewing: Deep kernel hedging · Jean-Loup Dupret, Donatien Hainaut and Edouard Motte · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Cost-Aware Deep-Kernel Hedging of Short Listed Equity Calls

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (7)

  • eq (3.1) self-financing: V_{t_{k+1}} − V_{t_k} = ξ_{t_k}(S_{t_{k+1}} − S_{t_k}) + η_{t_k}(B_{t_{k+1}} − B_{t_k}).: Deduct stock costs from the self-financing cash ledger and use subsequent adjusted-open stock fills. (invalidates: Frictionless paper Table 5 and Table 6 values as predictions for this test)
  • eq (3.2) terminal hedge value: V_T = B_T[v_0 + Σ_{k=0}^{n−1} ξ_{t_k}(S_{t_{k+1}}/B_{t_{k+1}} − S_{t_k}/B_{t_k})].: Use the same gain sum over actual stock-fill intervals, then deduct initial, subsequent and liquidation stock costs; the option is bought back at its observed earlier-exit quote. (invalidates: Frictionless paper Table 5 and Table 6 values as predictions for this test)
  • eq (3.3) gain and terminal shortfall: G^i_{t_{k+1}} := B^i_T(S^i_{t_{k+1}}/B^i_{t_{k+1}} − S^i_{t_k}/B^i_{t_k}); E_i(φ;v_0) := H^i − B^i_T v_0 − Σ_{k=0}^{n−1} φ(X^i_{t_k})G^i_{t_{k+1}}.: With B=1, build gains between real later adjusted-open fills; replace terminal H with the actual quoted option-exit liability and add stock costs to shortfall. (invalidates: Frictionless terminal-shortfall interpretation and numerical Table 5 and Table 6 results)
  • eqs (3.6)–(3.9) hedging representer: g_{θ,i}(·) := Σ_{k=0}^{n−1}G^i_{t_{k+1}}K_θ(X^i_{t_k},·); (Q_θ)_{ij} := Σ_{k,ℓ=0}^{n−1}G^i_{t_{k+1}}G^j_{t_{ℓ+1}}K_θ(X^i_{t_k},X^j_{t_ℓ}); min_φ J_θ(φ) = min_{α∈R^N}{(1/N)Σ_{i=1}^N L((H_{v_0}−Q_θα)_i)+λαᵀQ_θα}; φ*_{N,θ}(·)=Σ_{i=1}^N α*_{θ,i}g_{θ,i}(·).: Retain the gain-weighted RFF policy class, but optimize its coefficients directly with stock costs; do not claim the frictionless exact-RKHS representer objective remains equivalent. (invalidates: Exact frictionless representer equivalence for the cost-aware objective; the paper's exact-problem convergence claim for that changed objective)

3 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 22.0%Jul 2024
Sharpe
2.30
Total Return
22.0%
Max Drawdown
-2.8%
CAGR
4.5%
Volatility
2.1%
Beta vs SPY
0.00
Trades
4,859

What the paper reports for its own strategy

  • S&P500 out-of-sample (Jan 1 2023 to May 31 2026), ATM European call, quadratic hedging with order-3 signature features, frictionless (no transaction costs), P&L scaled by initial spot. Deep kernel hedging MSE / CVaR_0.95: 10 days 2.06e-5 / 1.00e-2; 20 days 3.13e-5 / 1.13e-2; 40 days 3.55e-5 / 1.36e-2 (mean over 10 seeds).
  • S&P500 out-of-sample (Jan 1 2023 to May 31 2026), ATM European call, CVaR hedging, frictionless. Deep kernel hedging CVaR_0.95: 10 days 9.19e-3, 20 days 1.08e-2, 40 days 1.15e-2.
  • S&P500 out-of-sample (Jan 1 2023 to May 31 2026), ATM Asian call, quadratic hedging with signature features, frictionless. Deep kernel hedging MSE: 10 days 1.30e-5, 20 days 1.76e-5, 40 days 2.96e-5.
  • Heston synthetic, ATM European call, quadratic, frictionless, 25,000 test paths, N=1000. Deep kernel hedging MSE 2.14e-5 (T=1/12), 6.88e-5 (T=1/4), 1.37e-4 (T=1/2).

A desk could keep the Heston gain and swap out the cosine features. The decisive comparison is in the appendix: an ATM European call on Heston paths, hedged under quadratic loss at T = 1/4 with N = 1000. A fixed layer of random ReLU units does almost as well as the Gaussian kernel features.

That finding changes how we read an otherwise strong set of results. Dupret, Hainaut and Motte beat a small deep-hedging network on simulated and index data under quadratic loss, and in most CVaR tests. Their learned embedding of the market state appears to do the work. Bochner's theorem is less central to the gain.

The hedger

At each rebalancing date, the hedge ratio depends on observable state: time, spot and variance, or a truncated path signature on real data. The authors place the hedge function in a reproducing kernel Hilbert space (RKHS) and apply a Gaussian kernel to a neural embedding. That embedding maps the state through a one-hidden-layer network with 16 tanh units to a 16-dimensional output. They normalise the output to unit length, then scale it by a learned 1/γ. The kernel bandwidth is learned along with the network.

Their theory gives a representer theorem for the hedging objective and, with a compact parameter set, existence of an optimal kernel and hedge. It also establishes almost-sure convergence of a random Fourier feature approximation to the exact problem. For the experiments, they draw D = 100 cosine features once and leave them fixed. Mini-batch AdamW jointly trains the embedding and a linear head over those features; under quadratic loss on Heston paths, it also trains the initial price v0. The objective is squared terminal hedging error or Rockafellar-Uryasev CVaR at 5%, with v0 fixed for CVaR.

The synthetic sample uses Heston with the Mikkilä-Kanniainen parameters (ρ = -0.79, vol-of-vol 0.62), 250 to 2,500 training paths and 25,000 test paths. The real-data sample uses daily S&P 500 observations, with January 2010 to December 2022 for training and January 2023 to May 2026 for testing. Its paths are overlapping rolling windows of 10, 20 or 40 trading days. On the ATM European call at T = 1/4 and N = 1000, out-of-sample MSE is 6.88e-5 for deep kernel hedging, 7.71e-5 for deep hedging and 1.07e-4 for practitioner Black-Scholes delta.

Why the raw kernel struggles on Heston paths

Remove the embedding and apply the RBF to raw (t, S, ν), and the plain Gaussian kernel loses to Black-Scholes delta in every tabulated European case on Heston paths. At N = 1000 and T = 1/12, its MSE is 3.94e-5 versus Black-Scholes' 2.56e-5. At T = 1/4, the figures are 1.33e-4 and 1.07e-4; at T = 1/2, they are 2.87e-4 and 2.40e-4. The authors say the classical kernel performs poorly and even underperforms the Black-Scholes benchmark in several settings. Their explanation is that raw input space lacks a suitable geometry. The real-data result differs: on S&P 500 paths with signature features, the plain kernel beats Black-Scholes on MSE at 10 days, 2.64e-5 to 3.12e-5.

A single bandwidth across three raw inputs treats a unit of time and a unit of variance alike, though their implications for delta differ. Learning the representation changes the Heston ranking. The embedded kernel gets the best number in the table.

What happens when cosine goes?

The authors' ablation holds the learned embedding in place and substitutes frozen random ReLU or tanh units for the cosine features. On Heston at T = 1/4 and N = 1000, cosine gives MSE of 6.87e-5 (the main results table reports 6.88e-5 for the same configuration), ReLU 6.90e-5 and tanh 6.88e-5. Take away the embedding instead, and all three trail Black-Scholes: 1.33e-4, 1.38e-4 and 1.45e-4.

The conclusion says the main benefit "lies in learning a data-dependent neural representation" and that the base kernel "appears to play a minor role". The abstract still credits "the inductive bias of kernel methods". Just before that conclusion, the authors argue that the edge over deep hedging "is especially relevant in low-data regimes, where the inductive bias inherited from the kernel appears to be particularly beneficial". The ablation, though, runs only at T = 1/4 and N = 1000. It does not compare cosine with random ReLU or tanh at N = 250, where the low-data claim matters most. Our reading that a frozen layer provides cheap regularisation for small samples is likewise untested at N = 250. With two other frozen random layers performing equally well, the choice of base kernel carries little weight in the tested setting. The paper also offers no test that keeps the embedding while removing the random layer. What remains is an embedding, a wide fixed random layer and a linear head.

Other details push our reading in the same direction. In our judgement, the explicit ridge penalty λ = 1e-6 is small beside the 5e-3 weight decay given to both models. The RKHS penalty therefore seems to add little beyond what the deep-hedging baseline already receives. The existence theory assumes a compact parameter set, which the authors disclose they do not impose. They also note that normalisation can break the injectivity needed by their universality argument.

These are two small networks in the empirical comparison. The deep-hedging benchmark is a 2x16 tanh FNN with 16d+305 trainable parameters; deep kernel hedging has 16d+389. The low-data gap is real. At T = 1/4 and N = 250, MSE is 7.91e-5 against 1.06e-4. It shrinks to 6.75e-5 against 7.24e-5 at N = 2500, and to 2.09e-5 against 2.21e-5 at T = 1/12. A learned embedding followed by a frozen 100-unit random layer looks like inexpensive regularisation when samples are small. That is useful to a desk, though narrower than the claim for a new kernel method.

Signature features on S&P 500 paths

With order-3 time-augmented signatures, deep kernel hedging has the lowest quadratic hedging error at every maturity on S&P 500 paths. At 40 days, normalised MSE is 3.55e-5, versus 4.35e-5 for deep hedging, 7.29e-5 for the plain kernel and 9.75e-5 for Black-Scholes. CVaR at 95% on that run is 1.36e-2, against 1.64e-2 for deep hedging and 1.76e-2 for Black-Scholes.

The choice of input moves the result further.

Give deep kernel hedging vanilla inputs (t, S/S0, realised vol), and its 40-day MSE rises to 8.93e-5. Deep hedging with signatures scores 4.35e-5. On that comparison, changing features matters more than changing hedger.

Under a CVaR objective, the ranking changes. Deep hedging has lower tail loss at 20 days, 1.04e-2 against 1.08e-2, and at 40 days, 1.10e-2 against 1.15e-2. Deep kernel hedging wins at 10 days, 9.19e-3 against 1.00e-2. The 40-day tail win comes with an MSE of 1.53e-4 for deep hedging, above Black-Scholes at 9.75e-5 and about twice deep kernel's 7.16e-5. The authors concede deep kernel hedging "does not systematically outperform deep hedging in terms of CVaR". They also caution that test CVaR draws on about 40 tail observations from 834 minus T test paths.

The paths overlap. 3262 minus T training paths come from one index history, leaving a much smaller effective sample than the count suggests. Dispersions are reported across ten seeds on a single split; they do not resample the data. We did not find market option prices in the real-data test. The payoff is hypothetical, v0 is a Black-Scholes price at realised vol, and the exercise is frictionless.

Our listed-options run

We built a traded version using end-of-day data from 2020-01-01 to 2024-07-01. It sells one near-the-money listed call with 25 to 45 days to expiry and strike/spot of 0.95 to 1.05. After 10 stock sessions, it buys the call back at the observed close. A signature-based deep-kernel RFF policy, trained on cost-inclusive CVaR of the shortfall, supplies the stock hedge. Hedge trades fill at the next open; notional is capped at 20%, with commissions of $0.004 a share and a $1 minimum.

The backtest returned 22.03% in total. Its Sharpe was 2.30, volatility 2.13% and maximum drawdown -2.85%. One-contract sizing explains the low volatility and shallow drawdown. Neither figure establishes hedge quality.

The nearest paper figures are the authors' CVaR-trained index results: 9.19e-3, 1.08e-2 and 1.15e-2 of spot at 10, 20 and 40 days. Their 9.19e-3 to 1.15e-2 CVaR of spot measures terminal hedging error, whereas our 22.03% is a total return. Their hedger starts at a Black-Scholes price calculated with realised volatility, without a market premium. We sell calls at observed prices, allowing any implied-over-realised premium to appear as profit. Our window includes the 2020 volatility spike and the 2022 drawdown; only about 18 months overlap their January 2023 to May 2026 test period. We also traded AAPL, AMZN, GOOGL, MSFT and NVDA. NVDA replaced the JPM we had specified, leaving attribution to our own design unresolved.

Our settlement code has intrinsic-floor and last-mark fallbacks. Missing quotes might trigger them, understate buyback cost and flatter returns. A profit factor of 11.98 alongside a 60.59% win rate makes that possibility worth checking: on a book short near-the-money calls, winners several times larger than losers fit the pattern understated buyback costs would produce. We cannot establish whether those fallbacks entered the reported metrics. This remains a hypothesis.

We ran no paired daily-delta comparator. We therefore cannot say whether the learned hedge beat delta or a deep-hedging network, the comparisons the paper makes, and we cannot fully explain the gap between the two sets of figures. This is one automated pass on our setup, not a verdict on the authors' work.

A paired test would change our view: the same listed contracts, fills and costs, comparing the learned hedge with quoted delta and a 2x16 FNN with a frozen random last layer. If the deep RFNN matched the deep kernel there too, the Fourier machinery could be dropped.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each eligible preselection close, choose the standard-deliverable call nearest ATM among the five named underlyings, with 25–45 calendar days to expiration and strike/spot in [0.95, 1.05].
Sell one selected contract at its next-session observed option close; commit to an option exit 10 stock sessions after entry.
At each decision close, construct causal normalized-spot and 20-day realized-variance paths, an order-3 time-augmented signature, and available option features.
Map the learned RFF share target through stock-notional, capital, leverage, and feasible-share limits; schedule stock hedge changes for the next adjusted open. Apply the same mapping and schedule to the logged delta comparator.
Precommit to liquidate stock at the option-exit-day open; buy back the option at its observed end-of-day close.
Train the learned policy on historical, normalized episode net-shortfall upper-tail CVaR, using feasible holdings and stock-trade costs rather than daily delta labels.

This is the specified procedure; the supplied code excerpt describes fallback settlement and corporate-action handling that differ from it, so the executed procedure needs reconciliation.