A delta-neutral straddle can leave a learned book persistently long Delta. Tan, Roberts and Zohren find a mean net position-normalized Delta of 0.127 in their LSTM trained solely to maximise Sharpe. This is net Delta per unit of gross position, so longs and shorts can offset. Even the first quartile is positive, at 0.035. Anyone running an end-to-end option-return strategy should check the same exposure in their own book this week. The reported Sharpe lift makes a less convincing case.
Building the straddle book
The authors trade one-month straddles on Nasdaq 100 names, opening a new trade at each monthly expiry. They pair the call and put at the strike nearest the money, within moneyness of 0.95 to 1.05, for expiry the following month. Initial deltas set the two leg weights so each straddle opens flat; those weights then remain fixed through expiry. OptionMetrics end-of-day quotes supply the data from January 2010 to December 2023, with positions marked at mid.
For each straddle, the network sets a daily position in [-1, 1], scaled by a 20-day EWMA volatility. Inputs include normalized returns over 1 to 20 days, multi-horizon MACD, moneyness, days to expiry and the five OptionMetrics Greeks. Option momentum follows Heston et al.: trailing average returns from past expired straddles over 1, 3, 6 and 12 months. The authors draw their return signal from trend and reversal patterns in straddle returns. Their four momentum benchmarks, TSMOM, MACD and the two Heston option-momentum rules, appear only as mean-reversion strategies, alongside a Long Only book. Momentum-based strategies, they write, "were generally unprofitable over the out-of-sample period".
The change to the LSTM is in its loss: negative Sharpe plus α times a Delta penalty. A penalty on raw Σ|XΔ| can be met by shrinking every position, an outcome the authors call signal shrinkage. Their exposure-normalized penalty (ENP) instead scales Delta exposure by gross position size. Their drift penalty (DP) scales it by Gamma exposure, charging straddles that have moved away from their strike. Both have L1 and L2 versions. The authors sweep α from 10^0 to 10^7, test with an expanding window in five-year increments, and aggregate several seeds.
At zero cost and a 15% vol target, the unpenalized LSTM earns Sharpe 2.131; the benchmarks range from 0.621 to 0.822. DP (L1) at α=10^2 is the best penalized run, reaching 2.339. Its drawdown is 0.210 versus 0.222 for the baseline, while mean net Delta drops to 0.056. The long tilt more than halves as Sharpe rises. These are separately trained models, and gross Delta barely changes. The narrow inference is that the higher Sharpe did not require the long tilt.
Does each position get hedged?
Very little. Gross position-normalized Delta averages absolute Deltas per unit of gross position, leaving no room for offsets. On DP (L1), its mean moves from 0.256 to 0.231 and its median from 0.272 to 0.238. The benchmarks run from 0.201 to 0.219, below all four headline penalized rows at 0.231 to 0.252. The authors acknowledge that the penalties work "primarily by reducing the baseline model's persistent positive bias" and only moderately lower gross Delta. Their conclusion calls the net-versus-gross distinction "A key contribution of this work" and says headline gross values are "largely similar to the benchmark strategies". The abstract gives a broader impression with "reducing realized directional exposure"; it leaves out that most of the reduction is net. Long and short Delta balance better across names, while individual straddles still drift with their underlyings.
More position-level hedging costs Sharpe. At α=10^4, DP (L1) cuts mean gross Delta by about 61% and earns 0.935. At α=10^3, the cut is about 41% and Sharpe is 1.753. Headline settings have much smaller effects: against the baseline's 0.256, mean gross Delta is 0.243 for ENP L1, 0.252 for ENP L2, 0.231 for DP L1 and 0.247 for DP L2. Those cuts run about 2% to 10%. The paper's stated 5% to 10% overstates the reductions for ENP (L2) and DP (L2).
The test-period choice of α
Only seven of 32 penalized runs beat the baseline's 2.131. The sweep comprises four variants at eight α values. The authors include α in a validation-tuned grid, as one hyperparameter in a 100-configuration random search. Yet Table 4 reports each α out of sample, and the paper says it shows "only the best performing model for each regularized variant". Every headline row is the highest out-of-sample Sharpe for its variant in that sweep. Selection therefore happens on the test period.
The DP (L1) path makes the risk of that choice plain. From α=10^0 to 10^4, its Sharpes read 2.139, 2.161, 2.339, 1.753 and 0.935. The winner stands just before a sharp drop. DP (L2) moves from 2.255 at α=10^0 to 0.759 at 10^4, then reaches -0.294 at 10^6. ENP (L2) varies less: from α=10^0 through 10^4, it stays between 2.007 and 2.260. Its peak of 2.260 exceeds the baseline's 2.131 by 0.129.
We did not find a seed count, seed dispersion or a significance test for the 0.208 Sharpe gap. A best-of-eight test-period pick without error bars is too thin to trade on that lift. The Delta finding has a different footing. Sharpe alone determined the four headline rows, yet their mean net Deltas are 0.056 to 0.101 against the baseline's 0.127. Nobody selected those rows for their net-Delta cuts.
Costs and the conclusion disagree
The paper charges costs per unit change in the vol-scaled signal. DP (L1) keeps Sharpe of 2.182 at 10 bps, 1.574 at 50 bps and 0.860 at 100 bps. Adding costs to the training loss ("TC Reg") raises the baseline at 100 bps from 0.816 to 0.871. For DP (L1), it lowers Sharpe from 0.860 to 0.777; at zero cost, that model falls from 2.339 to 1.803.
The conclusion calls L2 variants "slightly more resilient" to costs and says the best variant retains Sharpe of 1.053 at 100 bps. The table supports neither statement. At 100 bps, L1 leads for DP, 0.860 against 0.763, and ENP, 0.718 against 0.634. The section's own text says L1 "consistently" outperforms. The column's highest entry is 0.871. ENP with TC Reg at 100 bps is the one pair that favors L2, at 0.839 versus L1's 0.770. The DP pair runs the other way, 0.777 versus 0.690. No row approaches 1.053.
All results use mid prices and a proportional charge. Quote filters exclude zero bids, asks at or below the bid, and zero open interest; they impose no spread-width screen. Each trade crosses two option spreads. The reader is left to judge whether 100 bps per unit of turnover pays for them.
We are testing it now.
Our version fills at bid and ask and chooses α within each training window, rather than from the test sweep. Other differences are ours: a shorter period and an approximate large-cap universe in place of Nasdaq 100 constituents. For the paper, we did not find a statement saying whether membership is point-in-time. If validation-chosen α still brings net Delta toward DP (L1)'s 0.056 and leaves Sharpe above the unpenalized model after spreads, the penalty earns a production role. If the tilt reduction alone survives, learned straddle books still need net-Delta checks, and the 2.339 belongs in the drawer.