Entropic ambiguity at gamma-bar = 5 cuts the simulated American put from 0.11638 to 0.10881, a 6.5% reduction. The early-stop fraction moves much further, from about 0.41 to 0.82, across 1024 paths in a single trained run. A desk should care most about that asymmetry: exercise timing reacts far more sharply than price.

One limitation belongs up front. End-of-day data do not reveal the continuous-time Brownian dynamics, the ambiguity set or the BSDE driver. Nor can discrete EOD observations reproduce the continuous reflected quadratic solver. Any run we perform is therefore a daily discretized substitute, incapable of testing their scheme.

Ambiguity comes in two forms

Agram, Rems and Rosazza Gianin keep two sources of uncertainty separate. Under the usual model uncertainty, an adversary chooses a Girsanov change of measure, leaving the underlying drift unpinned. Discount-rate uncertainty follows the cash-subadditive setup of El Karoui and Ravanelli, with the adversary also choosing how future cash flows are discounted. Their general admissible set extends beyond a rectangle. In the worked examples of Section 6.4, however, constant rectangular bounds let the two knobs move independently. The authors call the object a paired-ambiguity risk measure, meaning a worst-case evaluation over pairs of measure and discount rate, penalised through the convex conjugate of a BSDE driver.

The first component of an upper reflected BSDE gives the optimal stopping value. Since the position is written as a liability, minus the payoff becomes the obstacle. Under their regularity assumptions, exercise occurs the first time the value process reaches that obstacle. In desk language, intrinsic value meets a continuation value penalised for discount and model ambiguity.

The entropic specification makes the problem tractable. For driver (gamma/2)|z|^2, an exponential transform converts the reflected quadratic BSDE into the lower Snell envelope of the transformed payoff e^{-gamma*xi}. Reflection remains after transformation, while the single conditional expectation from the unreflected case becomes a Snell envelope. The paper also establishes the expected structural properties for a risk measure on processes. It is decreasing monotone, cash-subadditive when the driver decreases in y, and concave when the driver is affine in (y,z). General convexity fails. A two-line deterministic counterexample uses xi_s = s and eta_s = -2s.

For the numerics, the authors use reflected deep backward dynamic programming. They truncate the driver at radius M_z and take an explicit min against the obstacle at every step. Their convergence argument joins the discrete-reflection and truncation estimates of Sun, Liang and Tang to the network error estimates of Huré, Pham and Warin. Three hidden layers of fifty neurons, tanh, Adam at 1e-3, batch 1024, fifty time steps. No market data enter the tests. The paths come from simulated GBM with S_0 = 1.0, K = 1.1, sigma = 0.2, r = 0.05, T = 1.0, alongside a driftless Brownian forward with a tanh-collared obstacle.

The solver's own 2.8% miss

The Black-Scholes case supplies the benchmark. A Cox-Ross-Rubinstein binomial tree with 2000 steps values the American put at 0.11973. Their solver produces 0.11638, leaving an absolute gap of 0.0034 and a reported relative deviation of -2.8%.

With entropic ambiguity switched on at gamma-bar = 5 and worst-case discount equal to r, the conservative price falls to 0.10881. The authors report -6.5% against their baseline of 0.11638. Their two prices differ by 0.0076, based on our subtraction, while the solver error is 0.0034. Comparing the ambiguity-adjusted mark directly with the binomial mark gives a total gap of 0.01092, also our subtraction. Roughly a third comes from the solver missing the vanilla benchmark. The contrast is stark: 0.41 to 0.82 beside a 6.5% price move.

Exercise timing is where the result lands.

In the Black-Scholes case, the early-stop fraction is about 0.41 and the mean conditional stopping time is 0.47. Entropic-discount ambiguity at gamma-bar = 5 raises the fraction to 0.82. Each figure rests on 1024 paths in a single trained run. The tested put begins 0.1 in the money at S_0 = 1, with an initial obstacle of -0.1.

The authors attach their own warnings to both numbers. The shift, they write, "is consistent with the increased contribution of the quadratic z term in the continuation region". They add that the experiment "does not by itself establish a comparative statics result for the optimal stopping boundary." They also trace the stopping-time cluster just below T to the imposed terminal contact Y_T = -xi_T. Some of both fractions is therefore mechanical rather than economic.

The usable range for the entropic radius is tight, by the authors' account. Their sweep reports 0.115, 0.115, 0.112, 0.109 and 0.107 at gamma-bar values of 0.5, 1, 2, 5 and 10. Below about 2, the effect is "buried in the optimal stopping bias". Above about 10, the price closes in on intrinsic 0.1 as the continuation region collapses. Their choice follows directly from that squeeze: gamma-bar = 5 "keeps the price clearly above intrinsic while the quadratic term contributes meaningfully." By our subtraction, the full sweep covers 0.008, about 2.4 times the 0.0034 benchmark gap.

Why gamma-bar = 5?

Nothing observable anchors gamma-bar, the discount band or the rectangular ambiguity set. The paper uses no listed-option surface and no realised exercise data. Its forward dynamics omit dividend yield, with drift r - sigma^2/2 and nothing else. No borrow. No transaction costs.

The paper is equally candid about the numerical machinery. Its convergence constant depends on the truncated driver's Lipschitz constant and therefore on M_z, which brings exponential terms. The authors call the estimate "qualitatively sharp but numerically weak." Exact truncation relies on a uniform bound for the discretely reflected integrand, requiring sigma to be deterministic and state-independent. The state-dependent volatility case remains open.

Two of the three test cases use the put payoff, which is nondifferentiable at the strike. As the numerics section acknowledges, the quantified rates in Proposition 5.4 consequently apply only to the collared tanh obstacle. We did not find an error-versus-step-size study anywhere. Despite the deep-learning framing, all three cases are one-dimensional. The sole dispersion measure is the stated 0.02 tolerance, described as twice the variation across training replicates.

The theory checks pass within that 0.02 tolerance. Monotonicity margins are +0.208, +0.212 and +0.214, while cash-subadditivity margins are +0.012, +0.008 and +0.006. Concavity registers +0.0044 for the affine driver, matching the prediction. The two non-affine drivers return -0.009 and +0.011. Both remain inside tolerance, so neither result resolves the question.

Limits of a daily implementation

We estimate the parameters from daily underlying histories, then examine them through sensitivity analysis. Our daily discretely reflected neural-BSDE implementation is compared with a discrete dynamic-programming baseline on US listed equity and ETF options. Because the options data are end-of-day, intraday exercise timing and intraday Greeks lie wholly outside what we can test.

The first application with figures is a valuation reserve for the tested contract type. The put starts 0.1 in the money at S_0 = 1. In their gamma-bar = 5 run, the ambiguity-adjusted mark sits 0.01092 below the binomial mark. Roughly a third of that gap is solver error. For the collared-obstacle example, we would also move the discount band, beta in [0.0, 0.1], to observe how rate shocks alter exercise policy.

We identified the same shape of problem in Brigo and Lucic on rough LSV filtering. The exponential transform gives a clean mapping from the reflected quadratic problem to a lower Snell envelope of e^{-gamma xi}. Whether anyone can use the result depends on a calibration step left for later.

A single out-of-sample exercise would change my view. Fit the model to a cross-section of listed American puts in one period, then score its implied exercise boundary against realised early exercise in the next, using the Black-Scholes boundary for comparison. Until that test exists, the 0.82 early-stop fraction remains a feature of a simulation with a hand-picked radius.