Liquidity regimes help a Hawkes model fit Dutch intraday orders, but the trade stream remains a poor fit. The Ljung-Box pass rate rises from 0.141 to 0.435, roughly tripling, while changing the ingredients of the regime index makes surprisingly little difference.
How the regimes work
Orders in continuous intraday electricity trading arrive in bursts that intensify toward gate closure. Coefficients of variation of interarrival times stay well above one across the session. Jhabli, AlSkaif and co-authors use a multivariate Hawkes process to model those bursts: an event lifts future event intensity, then its effect decays exponentially. Their four event types are added buy orders, added sell orders, buy trades and sell trades.
The model extends Morariu-Patrichi and Pakkanen's state-dependent Hawkes model with a composite regime variable and switching baselines. Its Liquidity Stress Index (LSI) averages three equally weighted z-scores: bid-ask spread, displayed book volume, and the 10-minute standard deviation of the raw mid-price. Each contract is divided at its own 33rd and 66th LSI percentiles into Low, Normal and Stressed states. Baselines, excitation amplitudes and decay rates vary by state. Penalized maximum likelihood fits 108 parameters per contract; the penalty pushes each state's branching-matrix spectral radius below 0.99.
The authors use EPEX SPOT order-book messages for Dutch XBID hourly contracts from January to March 2024. They fit only the last four hours before cross-zonal gate closure. The paper reports that 76% of transactions occur in the last four hours before delivery. Across 2161 contracts, the sample contains about 24.6 million events; added orders outnumber trades by more than 8:1.
Baselines increase with stress for every event type. Median stressed-to-low ratios run from 1.45x to 2.30x.
Endogeneity does not rise. The mean spectral radius is 0.893 in Stressed and 0.895 in Low (p=0.25), with Normal highest at 0.922. Trades excite added orders on both sides, at branching ratios of 0.26 to 0.40. Orders excite trades at 0.05 at most. Although the paper's text calls cross-side effects negligible, its branching matrices put opposite-side trade-to-order ratios at 0.31 to 0.40. The negligible cross-side ratios, 0.05 to 0.06, describe only the add-to-add channel. Under stress, buy-side trade self-excitation rises from 0.40 to 0.46. This is a descriptive study. We found no forecasting or trading exercise, and the paper does not claim one.
Do three states earn 108 parameters per contract?
The paper runs the appropriate single-regime benchmark. Its median Wasserstein-1 distance from the residuals to Exp(1) is 0.3707; the paper's index lowers that to 0.3080. Ljung-Box shows the larger improvement, with pass rates rising from 0.141 to 0.435. KS barely moves, from 0.314 to 0.344. Regimes chiefly soak up autocorrelation that a single parameter set leaves behind.
Three states beat two as well. Per-contract terciles yield W1 0.3079, versus 0.3536 and 0.3867 for global two-state cuts at the 75th and 90th percentiles. Since those cuts are global, global terciles provide the closer comparison: their W1 is 0.3170, still better than either two-state cut. States persist from event to event. The median probability of remaining in the same state is 0.96 for Low and 0.93 for Normal.
By our count, the benchmark uses 36 parameters against the state model's 108. We did not find an information criterion or held-out likelihood to charge the improvement for that extra complexity. The authors say these refits use a single optimization start and a reduced iteration budget, whereas the headline fits use three starts. Their shared protocol makes the refits comparable with each other. Their figures, including LB 0.435, cannot be compared with the headline fits. Single-start noise also shows up in the −V/+V gap below.
Spread makes the Wasserstein fit worse
The authors regard the index's insensitivity to its ingredients as a strength. Yet volume alone beats their three-component version. They call that index "close to the best"; it ranks 6th of 12 at W1 0.3080, while negative volume ranks first at 0.2934. Adding spread worsens W1 in every combination it joins. V+σ goes from 0.2964 to 0.3080, and −V+σ from 0.2946 to 0.3085. Ljung-Box points the other way: adding spread raises the pass rate in every pairing, including σ from 0.433 to 0.439 and V+σ from 0.426 to 0.435. Spread alone is the worst state-based variant, at 0.3549.
The −V and +V partitions are identical apart from their labels. Their 0.0013 W1 gap therefore measures the refit's noise floor; the top five compositions, from 0.2934 to 0.2964, fall within about twice that gap. Volume state means scarcely change, from 91.59 to 103.08, even though volume alone fits best. In January, median LSI also climbs through the session. We read the partition largely as a clock dividing the session, with stress explaining only part of it. Composition rankings remain stable across months (Kendall's τ 0.79 to 0.88), though we saw no standard errors for the W1 gaps. We found a similar single-component pattern before: volatility carried most of an order-book regime dial.
The trade stream
Added-order KS pass rates range from 71.7% to 81.4%. Trade pass rates are only 1.5% to 9.2%, and their W1 runs from 0.474 to 0.598. The authors acknowledge the mismatch and attribute it to sparse data: roughly 200 trades per type per contract in Stressed, or about 4 observations per transaction-source kernel parameter. In their view, richer kernels would worsen the problem.
The claim that stress mainly strengthens transaction self-excitation depends on this poorly fitting stream. Its branching ratio moves by 0.06, while the compensator fails KS in over 90% of contracts. The authors conclude that sparse transaction data cannot support reliable kernel estimation even with a single-exponential specification, yet their abstract leads with a transaction kernel result. Baselines and spectral radii receive paired Wilcoxon tests (p<10^-3 and p=0.25). The 0.40 to 0.46 transaction move is a cross-contract average reported without a test or interval.
Event construction may compound the issue. Trades include full (M) and partial (P) matches; one aggressor can produce several of them. In January 2024, the 25th percentile of trade interarrivals is 0.0001s. We did not find a step that combines a multi-fill sweep into one event. Without that step, sweeps would mechanically inflate trade self-excitation. A replication has to settle the event definition before fitting.
Decisions an implementer must make
- The z-scores use the mean and standard deviation "over the corresponding contract". This might mean the whole session shown in Fig. 5 or the four-hour fitting window summarized in Table I.
- Volatility is defined as a euro standard deviation of the raw mid, while the text reports it in basis points.
- Initial baselines use 0.5x the empirical rate for Low and 1.5x for Normal. We found no multiplier for Stressed.
- The authors run KS on 100 downsampled residuals per stream to reduce the test's sensitivity to large samples. At n=100, its power is limited.
- Cancellations (action code D) are never modelled as an event type.
The LSI has look-ahead by design. The paper calls it "an ex-post regime-classification device" and calculates z-scores and terciles from full-contract statistics. Live use would require trailing estimates, which the paper does not test.
We could not reproduce any of it. The model requires M7 message-level order books for Dutch XBID hourly contracts, with every add and fill alongside its book state. We only hold bar data.
The regime gain in serial dependence is the result we would trust, although the paper reports it only as an average across all event-state streams. A held-out likelihood comparison with trailing-window thresholds, in which three states still beat one, would make this worth building for execution near gate closure.