Sell the full 20 BTC at the first step. In the paper's reported training setup, that is the cheapest strategy tested, at 0.21 bps of implementation shortfall on the evaluation windows. TWAP costs 0.39 bps. Every learned configuration costs between 0.95 and 1.15 bps across those same 500 windows. The ranking is where a desk should begin.

Ardaiz, Budati and Habibnia acknowledge it in the abstract. They attribute it to "an environment whose frictionless replay and terminal-urgency penalty make early liquidation nearly costless". Their research question is seed sensitivity: apparent gains from an RL execution architecture may depend on lucky training seeds, and many runs are needed to tell.

The execution setup

The agent must liquidate 20 BTC (about $340k at January 2023 prices) through market orders over ten five-minute steps. The data comes from Kaggle: Binance BTC/USDT with ten levels on each side. Training covers January 9 to 14, 2023; evaluation covers January 16 to 20. Every column, down to prices and sizes at each level, is averaged into 5-minute bars. The agent's trades cannot move prices in this replay.

Child orders walk the averaged bid book. During the last two steps, that walk begins at the fifth-best bid, the authors' imposed late-liquidation penalty. Inventory left at the horizon is force-sold.

The tested architecture is a mixture of experts. It has K independent Double DQN agents, each built as a 5-layer, 30-unit network. A K-means router sends an episode to an expert according to features from the previous episode. The authors concede that the router cannot condition on regime: start bars are independently sampled, leaving those lagged features uninformative about the current episode. In practice, it divides the training data among experts.

They compare vanilla DDQL, K in {2, 4, 8}, and dense networks of width 64 and 93. Those dense networks match the parameter counts of K=4 and K=8. Each configuration has 100 training runs, evaluated on the same 500 out-of-sample episodes. The main result is null. At K=8, the paired mean difference from DDQL is -0.03 bps, with a 95% CI [-0.15, +0.09]. The authors also say equivalence has not been established.

Why does the agent sometimes sit out?

Training rewards slippage versus the prevailing mid; evaluation measures shortfall versus the arrival price. A policy that sits idle earns zero under the training reward until it incurs a single terminal penalty. Exploration never anneals in the reported configuration, either: epsilon ends around 0.88.

A collapsed policy therefore scores 1.47 bps. Eleven of 100 vanilla DDQL runs end there, along with 7 cap4 runs, 4 cap8 runs and 3 at K=2: 25 of 600 altogether. Neither K=4 nor K=8 has a collapsed run (Fisher p=0.0007 for each against DDQL). The authors identify suppression of this collapse as the one demonstrable effect of partitioning, then investigate what causes it.

Annealing changes the result

For that decomposition, the authors run on CPU and compare against a CPU baseline, keeping the CPU-versus-CUDA (GPU) difference out of the comparison. Annealed exploration alone moves DDQL from 12/100 collapsed runs to 0/100 (exact McNemar p=4.9×10^-4). Cap8 moves from 7/100 to 0/100 (p=0.016). Mean shortfall drops below TWAP for all three: 0.28 bps for DDQL, 0.01 for K=8 and 0.19 for cap8.

The 0.01 bps figure is no evidence of execution skill. The authors call it "execution at essentially the arrival price," possible only in the frictionless replay. They describe this modified setup as "a different diagnostic, not a better testbed."

Aligning training reward with the evaluation metric seems as though it should help. Alongside annealing, however, it sends DDQL collapse back to 19/30 (Fisher p=6.2×10^-6). Collapse rises to 48/100 with all three changes. The reward change by itself takes cap8 from a 3/30 baseline to 21/30 collapsed runs.

K=8 has no collapsed run under any of the six specifications. With all three changes, its mean is 0.15 bps against DDQL's 0.93. That repeated absence of collapse is the clearest stability pattern the paper connects to partitioning. We did not find a mechanism for it in the paper. The authors test replay retention as a candidate; changing the buffer alone has no detectable effect on DDQL collapse (5/30 against 2/30). I would put more weight on the K=8 pattern if it held in a simulator where the agent's own sales moved the book.

Seed spread falls; policy tails grow

Across-seed standard deviation of run-level mean shortfall goes 0.50, 0.45, 0.52, 0.37 from DDQL through K=8. The sequence does not improve steadily. K=8 clears Brown-Forsythe at p=0.034, then fails Holm adjustment (p_adj=0.17). With annealing, the gap shrinks to 0.42 against 0.44, while cap8 has the smallest spread at 0.32.

Risk within a trained policy goes in the other direction. Relative to DDQL, per-run CVaR-95 worsens by +5.4, +9.9 and +11.5 bps at K=2, 4 and 8. The K=8 comparison has p=0.006; K=4 has p=0.023, while the K=2 step has p=0.18. The authors see a possible explanation in thinner training data for each expert: at K=8, an expert receives roughly 312 of 2,500 episodes. They put the trade-off plainly: "Expert partitioning reduces across-seed dispersion at a real cost in within-policy tails." They later treat the dispersion finding as exploratory.

A trader holds one trained policy. The CVaR column consequently carries more weight than the across-seed spread. Its best learned configuration is dense, unpartitioned DDQL-cap8, at 104.1 versus DDQL's 110.8. The -6.6 bps gap is directional only (p=0.21).

How much does the replay charge for size?

Very little. The agent leaves no impact, and the visible book spans little price. Its ten levels have a median span of 0.60 bps; extrapolation beyond the bottom costs about 0.067 bps per level. Median visible depth is 20.9 units, although 46% of bars display less than the 20-unit order. The deep-book price floor never binds: three settings produce identical results in 90,000 evaluations. Algorithm 1 charges no fee. In this replay, selling a parent order larger than the displayed book is nearly free, which helps explain the 0.21 bps result alongside the 0.60 bps book span.

The out-of-sample evidence covers one five-day window. TWAP moves from -1.66 bps over the first 120 episodes to +0.39 across all 500. Confidence intervals capture seed variance conditional on that draw, nothing beyond it. Reproducibility poses another problem: switching training from CUDA to CPU changes 61% of paired evaluation values and shifts per-run means by as much as 1.58 bps. Campaign-level means nevertheless agree within 0.07 bps.

We could not rerun this. Our data consists of price bars; the policy needs the bid-ask spread and fills obtained by walking ten levels of depth. Without order book snapshots, we have neither the state nor the fills needed to reproduce it.

The authors' seed warning travels further than these execution costs. At 30 seeds, the annealing effect was not significant; at 100, it was decisive. An architecture comparison stopped at 30 seeds could have credited partitioning alone for fixing collapse and missed that annealing does the same.