The 19.2% tail-risk cut is worth a trader's attention, though it comes from a Heston market with steadily rising drift and long-run variance. Under uniform weights, far-future CVaR falls from 8.992 to 7.268. Once the authors charge proportional trading costs of 0.005, the joint method's cut is 6.6%.
How the training paths are challenged
Schneider, Looser, Garin, Liu and Kuhn begin with deep hedging, which learns a trading policy from simulated price paths. A small neural net at each trading date (two hidden layers of 20 units) turns the observed path into holdings. Training in the Heston experiments minimizes CVaR at 95% of hedging loss; the equity experiments use entropic risk. Their addition is adversarial training against a worst-case version of each batch, built around a recency-weighted training set.
The result they seek is a thinner loss tail when the next period differs from those paths. Far-future mean P&L is -4.528 for the plain hedge and -4.584 for the best one. The improvement lies in the tail.
The method has three parts. Recency weights make older training blocks count less, trading effective sample size against accumulated drift. One adversary shifts probability toward above-average losses within a KL budget τ. Another pushes each path along the direction that increases its loss, within a transport budget ε. Its metric takes the worst date for each variable. The authors name the combined attack WRAP (Wasserstein-Reweighting Adversarial Perturbation).
In their first-order expansion, reweighting charges for dispersion of losses across paths; transport charges for sensitivity to changes within a path. Their interaction enters at higher order. Training uses a one-step attack. Iterating it brings no improvement in the stationary test (3.545 against 3.543 test CVaR).
Does a linear trend flatter the 19.2% result?
The nonstationary experiment has eight simulated blocks of 4,000 paths. Drift rises from 0 to 0.20, and long-run variance from 0.04 to 0.16. Blocks 1 to 5 supply training paths, block 6 is for validation, and blocks 7 and 8 serve as the near and far tests. On block 8, the clean hedge's CVaR is 8.992. Reweighting alone cuts it 9.6%; transport alone cuts it 13.1%. Together they cut it 19.2%, to 7.268 (ε 0.30, τ 0.60). Recency weights bring the joint hedge down to 7.221.
The two attacks do complement each other. On the validation grid with uniform weights, joint budgets reduce CVaR 12.6%, versus at most 9.1% on either single axis (10.5% against 6.8% with recency weights). The inexpensive gain, though, comes from weighting. Without either adversary, it lowers far CVaR from 8.992 to 8.707 and turnover from 10.450 to 7.995, roughly a quarter less trading.
Budget selection is where I hesitate. Block 6 determines the budgets and the weighting ratio (0.00853, chosen from eleven candidates), while blocks 7 and 8 continue the same linear trend. Such a setup favors steady-trend extrapolation. The paper's equity data go the other way: test volatility is below validation volatility for all five stocks. The probability tilt is concentrated, too. At τ = 0.6, the effective sample size of the 4,000 validation paths falls to about 691. That gives the recent blocks considerable say over the inferred trend.
Trading costs change the size of the gain
The headline table excludes transaction costs. The authors also run the experiment with proportional costs of 0.005 and uniform weights. Far CVaR is 8.935 for the clean hedge, 8.434 for reweighting, 9.170 for transport alone and 8.349 for the joint method. The joint reduction is 6.6%. Transport alone makes matters worse.
These are the authors' figures, and they interpret them as support for joint training: it has the lowest risk both with costs and under a mean-CVaR objective. Under M-CVaR, far CVaR is 7.019 joint, 7.090 with reweighting alone and 7.251 clean. Still, the joint hedge gains about 1% over reweighting alone in both runs (8.434 to 8.349 in CVaR). Relative to clean, that cut has a price. Far mean P&L is -5.695 against -5.562, with costs of 1.130 against 1.002 per path. The paper notes that, relative to reweighting alone, the joint hedge is cheaper on both counts. At the same cost rate in the stationary market, its gain is 0.3% (4.494 to 4.481). Uniform reference weights are used throughout the cost section, leaving the turnover saving from recency weighting untested where it matters most.
Five stocks, generated paths
For the equity section, the authors fit generalized affine diffusions (volatility scaling as a power of price) to daily closes of AAPL, MSFT, AMZN, GOOGL and BRK-B. Fits for training use 2008 to 2019; validation uses 2020 to 2023; the test uses 2024 to 2025. The claim is a 30-step Asian call under entropic risk. Every test path comes from the 2024 to 2025 fits, with no option priced from the market. The authors describe this as a retrospective simulation benchmark conditional on fitted models.
The stock table gives a mixed result. Joint training beats the clean hedge on test in six of ten cases, ties in three when validation chooses zero budgets, and loses once (fixed-calibration BRK-B, 0.172 against 0.170). Several wins are within one seed standard deviation. The largest are GOOGL (0.339 to 0.302) and AAPL (0.288 to 0.267), both with resampled rolling parameters. Mean P&L falls in all six wins; turnover falls in all ten cases. The authors flag the calmer test window (BRK-B volatility 16.71% against 22.36% in training) and caution that the spreads reflect five training seeds only.
What would change my mind
A trading result would need daily listed option prices, spread and commission on every rebalance, chronological retraining on a rolling window, and ε and τ selected on a trailing validation slice. The authors put principled calibration of those two budgets in future work. Their conclusion also calls for direct training and evaluation on a large panel of realized market trajectories, without fitting a parametric model first. I would change my mind after a walk-forward on listed equity options where the joint hedge beats reweighting alone after costs. In the paper's cost run, that edge is about 1%.