The log-normal duration prior is the better order-flow model, though its point-forecast gain is thin. Replacing the geometric regime-length prior in Bayesian online changepoint detection wins every predictive-likelihood cell in Table 7.2, which covers 2022-08 to 2022-11 out of sample. For MSFT's August 2022 order flow, 11.7 likelihood units separate the models across 2739 observations. MSE slips from 0.8432 to 0.8422. Jebali gets the prior right. The size of the benefit remains the harder question.
The filter and its duration clock
Order-flow sign autocorrelation decays as a power law with exponent below one, leaving a divergent sum. Lillo and Farmer documented this independently of Bouchaud and co-authors. Order splitting supplies the standard explanation: a large metaorder reaches the market through a long sequence of child orders. Short-memory regimes with heterogeneous durations can reproduce that long-memory series, with each filtered regime standing in for a metaorder. Tsaknaki, Lillo and Mazzarisi used BOCPD for exactly this purpose, and their work is the report's starting point.
Adams and MacKay's filter uses the constant hazard 1/h, forcing geometric regime durations and therefore a characteristic timescale. Order flow lacks one. Jebali instead derives a run-length-dependent hazard H(r) from an explicit duration law in a hidden semi-Markov model, following Agudelo-España and co-authors. Pareto and log-normal versions are calibrated on NASDAQ data. The second half implements Knoblauch and Damoulas's BOCPDMS, placing Bayesian VARs inside the predictive model while two Inverse-Gamma hyperparameters are learned online by gradient ascent.
The data consist of LOBSTER trade prints for AAPL and MSFT from August to November 2022. Net signed volume per bucket is observed in volume clock, with every bucket containing N consecutive executions. In the calibration table, MSFT uses N=542, producing 2739 points for 2022-08, while AAPL uses N=887 and produces 2629 points. This tuned aggregation parameter varies by asset. August supplies the grid search, which is then carried into September, October and November. Out-of-sample therefore means three adjacent months for the same two tickers. Because volume-clock buckets for two names cannot be aligned in time, the multivariate analysis moves to 60-second calendar bars for June 2022 (T=8190).
Where does the log-normal win?
Every one of the sixteen predictive log-likelihood cells in Table 7.2 favors the log-normal prior across 2022-08 to 2022-11. MSE gives a less settled ranking. With the grid calibrated on MSE, log-normal wins 6 of 8 asset-months. Under log-likelihood calibration, Pareto wins 5 of 8, including all four MSFT months. Its values are 0.8675, 0.9162, 0.9183 and 0.8632, versus 0.9719, 1.1142, 1.0640 and 0.9618 for the constant-hazard baseline.
The abstract describes consistent log-normal dominance across assets, months and calibration criteria, attributing it to the absence of a characteristic timescale in order flow. The conclusion repeats the claim, covering the geometric baseline and Pareto alternative on both assets and under both calibration criteria. The Log-Lik columns of Table 7.2 support that wording. In the paper's global MSE summary for all out-of-sample datasets, however, Pareto is bolded in five cells within the log-likelihood-calibrated block.
The reversal says more than the ranking. On MSFT August, BOCPD's MSE jumps from 0.8432 to 0.9719, about 15%, solely because the grid-search criterion changes. Pareto moves from 0.8552 to 0.8675, about 1.4%. For MSFT, a duration-aware hazard makes the point forecast much less dependent on the criterion being optimised.
The reported MSE also needs an interpretive caveat. Definition 4.4 includes no normaliser, while table values cluster near 0.85 against MSFT bucket variance of 51.93. We read the univariate table as normalised by sample variance, matching the explicit definition in the multivariate evaluation section.
390 cuts through one month
With MSE calibration on MSFT August, the Pareto hazard declares 390 changepoints and a mean regime length of 7.0 buckets. Log-normal declares 66; the baseline declares 86. Their MSEs are separated by 1.5%. An order-of-magnitude change in segmentation therefore purchases 1.5% in prediction.
Jebali's conclusion places the primary blame for over-segmentation on the detection heuristic rather than the duration law. He redesigned that heuristic during the study after the original threshold rule generated spurious detections in synthetic data whenever prior spread was small relative to observation noise. In those cases, the most likely run length stays close to the axis. Pareto's sharp hazard peak just after d_min also matters, making the filter hyper-vigilant near the start of a regime. Anyone using these regimes beyond one-step forecasting should care that the hazard law alone changes the count by six times.
Appendix D makes the failure mode clearer. Across 1000 Monte Carlo runs with true hazard 1/100, the baseline filter detects average durations of 87.86 at prior spread 50 and 35.72 at prior spread 5. The theoretical mean is 100. Once regimes become difficult to identify, detection failure constrains the baseline. At wide prior spread, the duration-aware filters recover their assumed laws: Pareto detects 6.60 against 6.00 theoretical, while log-normal detects 10.65 against 8.24 at spread 50. Narrowing the spread produces the same sensitivity, with Pareto falling to 4.98 and log-normal to 9.88 at spread 5.
The bivariate extension falls behind
On June 2022 minute bars, the evaluation window runs from t_min=1170 through T=8190. A bivariate BOCPDMS using VAR(2) and diagonal noise records normalised MSE of 0.8707. Summed forecasts from two independent univariate BOCPDs reach 0.7705. Noise choice barely changes the outcome: among the five tested specifications, MSE ranges from 0.8707 for VAR(2) diagonal to 0.8916 for VAR(2) with the full observation covariance. Even an expanding-window OLS VAR(2), fitted without regimes, scores 0.8791.
The paper gives the mechanism. A BVAR coefficient estimated from a handful of observations within a short regime can be multiplied by a heavy-tailed volume spike. The resulting forecast may exceed the largest value ever taken by the data. A global OLS fit avoids that failure because it cannot adapt. Jebali's summary is direct: the adaptivity that helps BOCPDMS on regime-switching data becomes a liability when innovations are heavy-tailed and regimes are short.
The conclusion also concedes a structural limitation. Without an intercept, the BVAR cannot represent a per-regime level shift. Adding one would let the multivariate model nest the univariate model that beat it, since autoregressive coefficients can vanish and leave a constant mean within each regime. Per component, BOCPDMS VAR(2) scores 0.8582 for MSFT against 0.7351 from univariate BOCPD. An intercept preserves conjugacy trivially, and the conclusion says the implementation already supports a block-structured prior for it. It also says the model universe was never populated with competing specifications. The model-selection mechanism offered by BOCPDMS over BOCPD was therefore never exercised.
Ten bars control 7020
The 12% gap between the univariate pair and bivariate filter lacks statistical separation. Diebold-Mariano produces -0.78 (p=0.434) for MSFT and -0.48 (p=0.630) for AAPL. None of the pairwise differences in the paper's DM table reaches conventional significance, with the smallest p about 0.13. The same paragraph explains why: the ten worst bars among 7020 account for roughly 70% of total absolute loss differential. After winsorising the extreme 0.5%, the univariate advantage becomes DM 3.50 on MSFT and 8.60 on AAPL.
That tail concentration deserves more attention than the headline comparison. For this pair in this month, a handful of spikes make the DM test treat a 12% gap as a draw. Winsorise with disclosure, or use a loss that does not square.
We could not test any of these results ourselves. The method requires trade-by-trade signed volume with buyer or seller initiation from LOBSTER. We hold minute OHLCV, with no prints, no quotes and no book. Aggregation to OHLCV would replace the method's observable with a different one.
An intercept in the BVAR would change my view of the multivariate result. The current finding applies narrowly to a level-free VAR on two mega-caps over one month. The hazard result is more usable, though roughly one percent of MSE on two tickers improves the model without producing return.