The best network hedge has the same mean-variance error as an unrestricted hedge, though the theorem gives no finite network that attains it. The equality concerns an infimum: a gap can vanish across a sequence of strategies without vanishing for any strategy in that sequence.

Arandjelović and Schmock take up the deep hedging framework of Buehler et al. (2019), where a trading position is a measurable function of market information, represented by a neural network. They move the question to frictionless continuous time. Their algorithmic strategies are simple predictable processes. At each rebalancing time, a one-hidden-layer network with a bounded sigmoidal activation reads the time and finitely many stopped observations of a driving Lévy process Y. Its position stays in place until the next rebalance.

The proof is measure theory throughout. Starting with Hornik's 1991 universal approximation theorem, the authors extend its L^p result on Euclidean space to Orlicz spaces on general measurable spaces. A countable-subfamily reduction establishes that countably many stopped observations determine a variable known at a stopping time. Martingale convergence in Orlicz space cuts that countable collection down to a finite approximation; monotone-class arguments then extend the result to stochastic integrals. The paper contains no data or numerical experiment. Its main theorem approximates any stochastic integral against a semimartingale, in Émery's semimartingale topology, by integrals of algorithmic strategies. For X in H^p_sm with p in [1,∞), elementary algorithmic strategies are dense in L^p(X), with their integrals converging in H^p_sm. Three applications follow: mean-variance hedging, a no-free-lunch criterion over [0,1], and a semimartingale test using tanh-squashed positions bounded in [-1,1].

Why do the minimal errors agree?

Theorem 4.8 addresses a claim H in L^2, an asset X in H^2_sm, and a filtration G that may contain less information than the market's filtration. Terminal wealths produced by network strategies have the same L^2 closure as those produced by all G-predictable integrands. Taking the distance from H to either set therefore gives the same answer as taking it to their shared closure. The equality of minimal hedging errors is exact.

Attainment is another matter. The authors describe the network error as arbitrarily close to the minimum. In the martingale case, they exhibit network integrands converging to the optimal integrand U*_G in L^2(M), while the corresponding integrals converge in H^2. Doob's inequality, with constant 2, bounds the P&L error by twice the integrand error. An L^2(X) distance of at most δ from the optimum raises expected squared hedging error by at most δ².

Example 4.3 shows why that limiting language matters. For the Brownian family of targets tanh(∫h dW), held over (t,T], exact membership in the network class occurs only for a shy set of parameters h, the infinite-dimensional counterpart of measure zero. The empty-interior argument uses unit-normalized indicators separated by √2. Exact equality within this family is thus exceptional; approximation does the work.

The minimum may itself leave substantial risk. Remark 3.2 illustrates that a claim can resist replication even in the filtration generated by its own martingale. The example uses a one-jump process, which is not a Lévy process, with Z uniform on {-1, 0, 1}. No constant plus stochastic integral represents Z²: on {Z = 0}, any such representation misses by E[Z²] = 2/3. The abstract's unchanged-error claim says the network class adds nothing to the floor in Theorem 4.8. Equation (4.9) divides that floor into an unhedgeable residual and a loss from limited information. Neither receives a numerical value, and the paper gives no count of hidden units or observations needed to close the gap. The authors put the limit plainly: "the approximation theorems are qualitative."

A fixed daily grid

The elementary class E(ψ) trades at deterministic times, as a daily close strategy would. Such grids underpin the approximation for Theorem 4.8. Density arrives as the grids grow finer and the number of observations grows without a prescribed bound. A backtest fixes both.

Theorem 4.8 takes its infimum over every square-integrable G-predictable integrand, including continuously rebalanced ones. Put a network hedge and a delta hedge on the same once-a-day schedule and the theorem supplies no ranking between them. Implementation also leaves several choices open:

What can a backtest tell us?

We are running a test drawn from the paper's idea. A network uses information available at the prior close to predict a stock or ETF position. We will compare its out-of-sample hedging error with delta hedging on the same options. The results are not in yet.

Our run concerns a discrete-time neural hedge. It substitutes sampled trading for the paper's continuous-time setting, whose convergence limits sampled market data cannot reach. No backtest of ours can confirm or refute those convergence claims.

The authors also acknowledge the absence of frictions. Their conclusion leaves portfolio constraints and transaction costs for future work, although Buehler et al. (2019) include costs in their discrete-time framework. Those choices remain with the implementer. If a network beats daily delta out-of-sample after costs, the finding will depend on the data and training procedure. The one-hidden-layer theorem licenses the search; it offers no forecast of the result.

The paper has a further theoretical application through Bichteler-Dellacherie. In a filtration generated by a Lévy process, a càdlàg adapted process is a semimartingale on [0,T] exactly when terminal integrals of algorithmic strategies, squashed by tanh into [-1,1], are bounded in probability. For a trader, the paper supports trying network hedges while leaving their realized performance open.