Keep the closed-form hedge if you have one. Wong and Campajola show how far PPO can stray from a linear-quadratic answer, then trace the failure to what its critic sees in nearby actions.
The game has an answer key
The broker-trader game comes from Cartea, Jaimungal and Sánchez-Betancourt. A broker handles orders from an informed trader, who acts on a mean-reverting private signal, and uninformed clients, whose random flow also mean-reverts. The broker hedges its inventory in a lit market, paying quadratic execution costs. Its trades also move the price through an impact term. With zero resilience in every run (p = 0), that impact persists. The equilibrium hedge is linear in broker inventory, informed inventory, impact, signal and uninformed flow; Riccati equations determine its time-varying coefficients.
The authors give PPO control of the broker. When PPO deviates, the informed trader still follows its analytical feedback rule. The opponent therefore is not playing Nash, as the authors acknowledge. They discretise at Δt = 2×10⁻⁴ (5,000 steps), with second-order Δt² corrections in the per-step reward. A finite-grid Riccati recursion gives the reference action for each state PPO visits. It also gives the "LQ loss": a curvature-weighted squared action error equal to the exact expected payoff gap.
All experiments are synthetic. Each Section 4 configuration uses five seeds, 4 million transitions per seed and 250 test paths. The residual experiments use five seeds, 100 critic warm-up plus 800 joint updates, and 100 test paths. Earlier RL execution work adjusted an Almgren-Chriss schedule or measured liquidation curves against a benchmark. Here, each poor action can be checked against the right action in the same state. The simulator receives a check too: against the continuous solution, its training grid has action RMSE 0.003389 and LQ loss 1.383×10⁻⁸.
What changes when client flow becomes stochastic?
Switch off uninformed flow and the feedforward PPO actor comes close: LQ loss is 0.045 ± 0.031, with action RMSE 5.494 ± 2.163. This run starts with wider exploration (log std 3); the others use 0. The LSTM reaches 16.304 RMSE.
Raise flow volatility to 100 and the gap opens. The clipped reference earns 0.1303, while PPO-FFNN earns −7.060 with RMSE 20.449. PPO-LSTM earns −39.194 with RMSE 31.284. Actions are clipped to ±100. On a typical step, the FFNN error spans a fifth of that half-range.
There is a table discrepancy worth keeping in view. For the same nominal full-information, ReLU FFNN configuration at flow volatility 100, the ablation table reports −5.75 ± 4.82; the summary table gives −7.060. We did not find the two reconciled. Either figure supports the same conclusion.
The critic's ranking problem
Capacity and reward implementation look unlikely to explain the FFNN result. With uninformed flow at zero, supervised training on the reference action produces FFNN test RMSE 2.31, rising only to 2.06 when it drives the simulator. The supervised LSTM fits nearly as well open-loop, at 2.62. Once it controls the simulator, RMSE runs from 11.85 to 339.70 across seeds. Open-loop fit is a poor assurance for that recurrent controller. The supervised check was not repeated at flow volatility 100.
Two reward ablations barely move payoff. Removing the Δt² term changes it by +1.48 [−7.34, 10.31]; removing the midprice input changes it by −1.09 [−8.56, 6.37]. Multiply the reward by 100 and payoff changes by −29.24 [−33.93, −24.54], falling on every seed.
The critic explains 0.93 of return variance across states, yet struggles with the choice PPO needs it to make. Given two actions 20 units apart, it picks the better one in 51.7% of 120 test states. Pearson is 0.216; Spearman is 0.053. Among the 71 states where Monte Carlo clearly distinguishes the actions, accuracy drops to 49.3%.
A coin flip.
The authors call this "a plausible source" of failure, and that restraint fits the evidence. Potential-based shaping with the exact LQ value function should give PPO a clearer local signal. Payoff moves by +1.57 [−5.00, 8.14]. Three of five seeds improve, while per-step reward volatility rises about 75%. I take that as evidence that the critic problem survives even a perfect hint. A different actor-critic passing the same ranking test would change my view of how general the failure is.
History gives the filter enough
Under partial information, the broker cannot see informed inventory or the signal. The certainty-equivalent controller uses its own inventory change to recover informed flow, then sums that flow into an inventory estimate. It inverts the informed trader's linear rule to estimate the signal, reaching RMSE 0.0141. Its payoff is 0.1297, just 0.00061 below the full-information reference. Partial-information PPO earns −21.680 (FFNN) and −24.108 (LSTM).
This inversion depends on the simulator's informed trader following a known linear rule with known coefficients. A real client's rule is unknown, leaving room for a learned estimator. PPO never reaches that point here.
A small gain from residual PPO
When execution cost changes, the broker retains the policy solved at cost 0.0012. PPO learns an additive correction without observing the new cost. At 0.0006, the reference earns 0.53010, the frozen policy 0.29632 and the adjusted policy 0.30151. The gain is 0.00519 [0.00296, 0.00743], across all five seeds. That closes 2.22% of the payoff gap. The authors describe the gain as small and repeatable; both descriptions hold. At 0.0012 and 0.0018, the change is 0.00019 both times. Their 95% intervals, [−0.00041, 0.00079] and [−0.00034, 0.00071], include zero.
I part company with the conclusion's description of the frozen policy as an "accurate starting strategy". The authors have a case for why corrections remain small: wider exploration, using the ReLU setup from the full-information runs, reduced payoff on all five residual policies. They read that as exploration pulling the adjustment away from a useful starting point. Yet at halved cost, the frozen policy earns 56% of the reference. The largest residual action on the halved-cost test paths is 1.41, and final-action clipping affects 0.029% of decisions. The warm start keeps PPO near a policy leaving 0.234 on the table. A desk able to estimate the new cost could recompute the Riccati coefficients, as the reference does.
Why we built no residual-broker backtest
The broker needs informed and uninformed client order flow as well as its own inventory. Its decision grid runs 5,000 steps per horizon. Minute equity bars contain none of that flow. Running an inventory schedule on those bars would fail to reproduce either the game or its comparison.
The ranking test is the result I would carry onto a desk: before trusting a PPO execution agent facing stochastic flow, see whether its critic can distinguish nearby actions. Here, it could not.