A five-asset uncertain-volatility pricer can hit the second decimal and still give you a poor Gamma. The price interval exposes the problem: the primal-to-residual-dual interval is more than six times the PINN's, 0.217 against 0.034. Abbas-Turki, Chassagneux, Lemor, Loeper and Sananes make that discrepancy the useful result. On their geometric call spread, with a reference of 9.70, all four learned candidates start between 9.6942 and 9.7074. Their residual-based upper estimates spread from 9.7248 to 9.8991.

How the dual uses Gamma

In the uncertain volatility model, an adversary chooses volatilities at every instant and, in several dimensions, correlations within a box. The seller's value solves the Black-Scholes-Barenblatt (BSB) equation. Its Hamiltonian H* takes the supremum of Tr[Γaaᵀ] over admissible covariances; Γ is half the cash-gamma matrix. In one dimension, the choice is σ_max for positive Gamma and σ_min for negative Gamma.

The authors begin with any smooth candidate v. Its Gamma picks a feedback control. Simulating the payoff under that admissible control gives a lower bound. For an upper bound, their dual fixes a Gamma field and maximizes over controls the payoff plus an integrated gap: H* less the Tr[Γααᵀ] generated by the control actually played. The gap is nonnegative, so the construction dominates the price for every field and equals it at the true Gamma. Using the candidate's Gamma produces U_D. The equivalent U_ε starts with the candidate value, then adds a control problem involving the BSB residual and terminal mismatch.

The practical controls are piecewise constant across N steps. Restricting them this way costs at most C(T/N)^{1/4}, using a rate the paper takes from Jakobsen, Picarelli and Reisinger. Candidate values come from backward stochastic policy gradient (SPG) critics, with one network per date. Training uses values alone, gradient matching (Sobolev-1), or gradient and random directional second-derivative matching (Sobolev-2). The other candidates are a PINN, a tanh network trained against the BSB residual, and a finite-difference (FD) solution in two dimensions. The comparison is the gradient-only martingale dual of Henry-Labordère, Litterer and Ren. All three payoffs are synthetic.

The interval follows the Hessian

Both Hamiltonian duals use Γ directly. Hessian noise therefore enters at every step along every path. Against Sobolev-1, Sobolev-2 moves five-asset U_ε from 9.7893 to 9.7526 and U_D from 9.7618 to 9.7329. The martingale dual goes from 10.2176 to 10.2468 instead; its estimator never uses curvature. On the best-of butterfly, whose reference is 6.70, U_D is 7.9733 for the plain critic, 6.9811 for Sobolev-2, 6.7247 for the PINN and 6.6997 for FD.

Poor candidates show the cost more starkly. For the outperformer spread, with a reference of 12.83, the plain critic gives U_ε of 15.5486 and U_D of 14.6497. Both exceed the martingale estimate of 14.0547. Its primal price is 12.7710, within the 12.7710 to 12.8008 primal range across all learned candidates.

Price alone tells you almost nothing here.

Rebuilding the candidates

The SPG training details are fairly specific: 2^16 states per date, mini-batches of 2^11, one GELU layer of width 32, and Sobolev weights η1 = 1 and η2 = 10^-2. The PINN uses three tanh layers of width 64, 2^8 interior and 2^8 terminal points per step, λ_T = 10, and 10^4 Adam iterations. Resimulation uses 2^19 paths. A rebuild still needs the authors' companion SPG paper for the actor parametrization, PPO update and state sampling.

Some choices remain with the implementer. We did not find a separate N for the butterfly, only a statement that it reuses the outperformer methodology (N = 128). The paper calls H* explicit for the correlated 2D box but writes out only the zero-correlation five-asset formula, leaving the 2D maximizer over σ and ρ to be derived. The martingale dual also uses 2^10 outer paths, against 2^19 elsewhere, so its intervals are wider than the Hamiltonian intervals by construction. For the plain outperformer critic, its interval is ±0.0524, compared with ±0.0199 for U_D and ±0.0071 for U_ε on that same critic. We did not see seed-to-seed variation reported for any candidate.

Why the upper estimates remain uncertified

The authors explicitly say the reported duals are not guaranteed upper bounds. Replacing each supremum with an SPG-learned control can only lower its value. As they put it, "the loss of the upper-bound property therefore comes from the optimization error rather than from the Monte Carlo approximation." Their intervals account for sampling error only. For time-grid SPG critics, reported U_ε also omits a consistency defect with no rate and the (T/N)^{1/4} term. At N = 64, that factor is 0.354. The paper supplies no value for C, leaving the correction's size unknown. The gap available to absorb it is small: 0.014 between the PINN's 5D U_D (9.7143) and the 9.70 reference.

The FD rows put U_D at 6.6997 ± 0.0130 and 12.8291 ± 0.0189. Both point estimates fall below their references, 6.70 and 12.83, by less than the Monte Carlo half-widths. Those references are themselves uncertified FD values. The rows therefore cannot distinguish a valid bound from a downward-biased estimate. The authors treat them as a consistency check on the two implementations and describe the construction as bracketing the price "up to the numerical errors introduced when these quantities are themselves approximated." My further inference is narrower: optimization bias only lowers a resimulated dual, so a dual far above the primal still flags a bad candidate. A loose U_D is evidence against the pricer. A tight U_D remains an estimate.

The cost of derivative supervision

Derivative supervision roughly doubles training time, an additional cost the paper calls moderate. On the outperformer, Sobolev-1 takes 223s to train and Sobolev-2 takes 294s, versus 148s for the plain critic. Each dual evaluation adds 249 to 291s. In five dimensions, training rises from 87s to 175s.

For a team checking a UVM pricer with a good candidate, whether FD, PINN or a Sobolev critic, U_D was the tighter diagnostic in these experiments. With the FD candidate, the martingale estimate reads 13.2104 on the outperformer and 7.4190 on the butterfly. U_D comes within 0.001 of both references: 12.8291 vs 12.83 and 6.6997 vs 6.70. The plain critic reverses that ordering on both 2D problems. Its U_D is 14.6497 against U_M of 14.0547 on the outperformer, and 7.9733 against 7.8710 on the butterfly. The authors call the ordering "an empirical observation rather than a general ordering between the different dual formulations."

These experiments say nothing about hedging P&L or market data. The five-asset case fixes correlations at zero, so the elliptope constraint, which binds only for d ≥ 3, is inactive throughout. I would treat U_D as a bound after a harder control search lifts the FD U_D and it stays within 0.02 of 12.83.