Force one learned liquidator into a profitable first trade, and its rival sells early enough to wipe out the gain within four periods. The rival sees only prices, never the other agent's orders. Koulouris and Campajola find this response in all ten trained pairs. Lillo and Macrì had identified controlled deviations as the next test; this paper performs one. The result comes from a ten-period, two-player simulation with almost no price noise. Its reach into live markets remains untested.
The liquidation game
The agents are neural-network traders trained by proximal policy optimisation (PPO) to minimise their own implementation shortfall. They play a discrete Almgren-Chriss liquidation game. Each starts with 100 units of one asset priced at 10, must finish selling over 10 periods, and trades at a shared execution price determined by both orders. Temporary impact is a = 0.002; permanent impact is kappa = 0.001. The agents are risk-neutral and have no shared reward.
The competitive reference is the closed-loop Nash equilibrium. At the cooperative reference, joint TWAP, each agent sells 10 a period. Selling slowly together reduces impact for both, yet either agent can lower its own cost by leaving that schedule while the other sticks to it.
Each independent PPO actor-critic takes continuous actions. A small Transformer (2 layers, 2 heads) reads prices and the agent's own trades within the episode. Opponent inventory and orders stay hidden; the price path is its only clue. Ten pairs train for 40,000 episodes each, then face 500 frozen test episodes.
In 10 of 10 runs, both agents' costs fall below Nash and remain above joint TWAP. The authors call the outcome "supra-competitive, while falling short of the cooperative benchmark." On average, the agents sell less than Nash early and leave more liquidation for later. Lillo and Macrì had already reported below-Nash costs from Double Deep Q-learners. The same two authors' earlier paper linked the effect to access to history. All these shortfalls are simulated under linear impact. No fill was observed anywhere.
What does a forced first trade provoke?
The authors first train a fresh learner against a non-adaptive player that follows the pooled mean schedule. In all 10 runs, the learner does better than copying it. Its profitable deviation, d, establishes that pairing the mean schedule with itself is no equilibrium. The authors then take the original frozen pairs and force only d's first trade, the first-trade deviation (d1), onto one agent. That agent resumes its own policy from step 2. Its opponent follows its learned policy throughout. They repeat the test with the roles swapped.
In all 10 run means, the opponent, or punisher, sells more in timesteps 2 to 5 and less thereafter. Those early sales push down the prices received on the deviator's remaining inventory. G denotes the gain from deviating if the opponent does not react; H is the extra cost caused by its reaction. Across all 10,000 paired cases, G > 0 and H > G. The trade pays if left alone and loses money once the opponent responds.
A second check uses a total-variation bound from the impact model. It tests whether shifting enough of the punisher's 100 shares through time could account for a loss of that size. The bound passes in every run.
The punisher's own cost
The punisher's average payoff is "materially unchanged" against its path without punishment under the same forced deviation, according to the paper. We did not find a number for this in the text; the comparison must be read from Figure 4. The authors also sample "thousands of nearby feasible schedules" for the punisher. Some pay it more while doing less harm to the deviator.
That comparison carries the sharper evidence of punishment. Against the unpunished path, the learned response costs the punisher essentially nothing. Against some sampled alternatives, it gives up profit to hurt the deviator. The sampling is loosely specified: the text gives no count of schedules that beat the learned response. Figure 4 presents per-run mean differences without numbers in the prose, and the paper does not compare the response with a Nash best reply to the deviation.
One deviation at sigma 1e-9
The authors keep the claim scoped. The abstract says both checks hold "for the tested deviation". Their conclusion asks for stronger price noise, larger populations and a broader range of deviations. Immediately before that, it says the findings support a collusive interpretation "beyond the observation of supra-competitive costs alone." We agree for this game. The agents use history to erase a deviation's profit and pass up better sampled schedules in the process.
The apparent 10,000 cases offer less independent evidence than the count suggests. Section 3 says the 500 deterministic test episodes per run produce "essentially overlapping cost pairs at the figure's scale" for undeviated costs. We infer much the same for the deviation tests: action selection is deterministic, and sigma is 1e-9. The independent units are the 10 trained pairs, tested with either agent as deviator. Ten of ten pairs pass in both roles. That remains a small count.
The training logs make a weaker case. Spontaneous first trades within 0.01 of d's first trade peak around episode 18,000, then decline. The authors interpret the decline as deterrence. Exploration log-std, however, falls from -1.2 to -4 over the first 75% of training, or 30,000 episodes. A narrowing policy could also yield fewer matches. "Consistent with" deterrence is the authors' phrasing, and the evidence supports no stronger reading.
A clear signal at step 2
Volatility is sigma = 1e-9, so the price at step 2 reveals the deviation exactly. Detection through prices is central to the response. The authors' proposed test with stronger price noise would show whether the punisher still spots a one-trade departure when noise approaches the size of impact. The game also has one asset, two symmetric 100-share inventories, linear impact and no risk penalty. We did not reproduce the simulation.
Koulouris and Campajola move beyond showing that RL execution agents beat Nash. They bring the pricing-game standard of Calvano et al. into execution through an imposed deviation, a paired counterfactual and an accounting of both payoffs. The next test is the same intervention with real noise. If retaliation survives a sigma large enough to blur the step-2 signal, the finding begins to matter to anyone running two learned execution algos on the same name.