A trader should care about a smaller calibration grid if it preserves the fit. Afzali, Della Corte and Papapantoleon find exactly that on synthetic rough Heston surfaces: a network given about a third of the long-maturity grid calibrates as well as one given the whole grid.
Heston and rough Heston model stochastic volatility. Their parameters are κ (mean reversion), γ (vol-of-vol), ρ (spot-vol correlation), v0 (initial variance) and v̄ (long-run variance). Rough Heston adds a Hurst exponent H; variance paths grow rough when H sits well below 1/2. Fitting these models to an implied volatility surface normally takes repeated optimization. The authors train networks offline to do the calibration in one forward pass, then use SHAP and νSHAP to examine which parts of the surface inform each parameter.
Short maturities and smile wings carry the strongest attributions. On the rough Heston long grid, the MLP and spHW maps look broadly alike. A νSHAP-guided cut removes about 67% of that grid's inputs without hurting calibration accuracy.
The authors call their interpretations "stated as qualitative evidence rather than as identification results." The abstract makes a stronger claim: the SHAP and νSHAP differences "expose substantial redundancy in the calibration input," and νSHAP can "guide a significant reduction in input dimensionality." The sufficiency calculation below puts pressure on that reading of redundancy. None of these results yet establishes a volatility signal anyone can trade.
How the network learns calibration
The network learns to invert the Heston pricing map. Latin Hypercube Sampling supplies parameter vectors, which are priced into implied volatility surfaces. Heston pricing uses COS. Rough Heston solves the fractional Riccati equation with a fractional Adams scheme before applying COS. Training then teaches the network to recover parameters from a surface in one forward pass.
The sample contains 2·10^5 Heston surfaces on an 11×8 grid (88 inputs) and 10^4 rough Heston surfaces per grid. Rough Heston has two grids: a long grid of 9×6 = 54 inputs spanning 0.60 to 2.00 years, and a short grid of 11×10 = 110 inputs spanning 0.02 to 0.50 years. Inputs are MinMax-scaled and then ZCA-whitened. ZCA is a decorrelating linear transform designed to stay as close as possible to the original coordinates.
The authors compare three network families: MLPs, highway networks, and a softmax two-gate highway (spHW). They prove that spHW spans the same function class as a standard highway layer. Its average parameter MAE is lowest on every grid: 0.0079 for Heston, 4.6e-4 on the long grid and 0.0027 on the short grid. Model size tells a less tidy story. On the long grid, a 1,526-parameter MLP reaches 0.0010, against 0.0021 for a 27,526-parameter highway net.
SHAP measures the change in a prediction when a feature is pinned to its observed value while the rest are drawn from background data. νSHAP instead asks whether a pinned subset keeps the prediction within tolerance; each feature gets credit when it completes a sufficient set. The authors compute both for 50 test surfaces against 100 validation surfaces. Short maturities and smile wings dominate the resulting maps. SHAP concentrates its attributions, while νSHAP spreads them more widely. The authors interpret that spread as surface redundancy.
The long grid cut
Here is the engineering result. Keeping the short-maturity region of the long grid removes roughly 67% of the inputs. After retraining, spHW average MAE falls from 4.6e-4 to 3.9e-4. Mean repriced-surface RMSE drops from 0.041 to 0.026 vol points, and the median drops from 0.031 to 0.022.
The gains are uneven. κ stays at 0.0011 and v̄ at 3.6e-4. γ, v0 and H improve, consistent with the paper's link between those parameters and short maturities; ρ improves too. The paper expects κ and v̄ to matter more at relatively longer maturities, yet removing those maturities leaves both errors unchanged.
The authors allow that the improvement "may also partly reflect training variability." We found no repeated seeds or confidence intervals to put an error bar on the 0.7e-4 average MAE gap. Explanations of 50 test surfaces also informed the retained region, and the reduced model was scored on that test split. The leak is small and built into the procedure.
A practitioner could have picked the same region without νSHAP: it is the front of the grid. The νSHAP selection therefore cannot be distinguished here from that rule of thumb. We did not find a random-subset or PCA-truncation baseline either. And "short" starts at 0.60 years on this grid.
The parameter priors leave little room. H runs from 0.1286 to 0.1766; ρ runs from -0.7071 to -0.5940. Relative to the H window, the reduced-grid spHW error of 1.7e-4 still amounts to about 0.35%, so the accuracy survives that scaling. The network nevertheless never sees H outside 0.13 to 0.18 or ρ outside -0.71 to -0.59.
Which quotes are being discarded?
The authors warn against treating the heatmaps as pointwise importance scores for individual strike-maturity contracts. Whitening makes each coordinate a linear combination of all 54 vols. Yet the reduction discards actual grid points selected from those maps. ZCA stays close to the original basis, which gives the choice a reasonable footing. It remains a judgment call.
A thin background for νSHAP
The paper acknowledges a problem with its sufficiency test: when no background surface lies within tolerance on a coalition, sufficiency holds vacuously. Tolerance is ±0.05 per whitened coordinate. Our arithmetic is an upper-bound estimate. Suppose whitening yields roughly unit variance and the coordinates can be treated as independent, though whitening removes correlation rather than guaranteeing independence. The band then contains at most about 4% of a coordinate's mass, nearer 3% for a typical surface. Among 100 background surfaces, one feature finds three or four matches; a pair finds about 0.1, almost none.
Most coalitions of three or more features would consequently pass by default. Such passes would flatten the contrast and make νSHAP maps look diffuse. Since the paper takes diffuseness as evidence of redundancy, diffuseness alone cannot establish that claim here. The short-maturity tilt used to select the smaller grid must survive in the singleton and small-coalition stage. A reported share of coalitions with at least one background match would resolve this concern. A high share would change my view.
Parameter error versus surface fit on the short grid
On the rough Heston short grid, spHW records the lowest average MAE, 0.0027, while posting the worst mean repricing RMSE: 2.40 vol points versus 1.20 for the MLP. Their medians are close, 0.18 versus 0.15. The authors read the large means alongside medians of 0.15 to 0.18 for all three networks as evidence of a skewed error distribution with occasional poorly reconstructed surfaces. They do not explain why spHW's mean doubles the MLP's.
Those occasional surfaces matter to a trader.
What remains before a trade
All the paper's results use synthetic surfaces, as the authors state. It prices synthetic European calls; no traded market enters the study. Our own exercise uses listed US equity options, limited to calls on non-dividend-paying stocks. That restriction narrows the European-American gap without closing it. Our exercise is still running, has no numbers yet, and is neither a replication nor a test of the paper. Synthetic-surface results cannot be assumed to carry over to market surfaces.
For reduced-grid repricing gaps to carry a relative-value interpretation, the networks would first need to calibrate end-of-day market surfaces with bid-ask noise and missing wings. Full-grid and reduced-grid parameters would then need a day-to-day stability comparison on dates outside training. We found no runtime comparison with an optimizer, leaving the speed claim asserted rather than measured.
A 67% smaller input without lost accuracy is worth having. The paper demonstrates it on 10^4 synthetic surfaces, for one grid and one architecture.