For an equity desk, the event-model edge is thin: 0.009 macro-F1 on AAPL and 0.002 on BAC at h=10. Schaurecker, Strand, O'Sullivan and Jakob find a larger 0.038 gap in intraday electricity and an equity edge of 0.002 to 0.009. Their engineering case deserves attention even where the classification gain is small. Before scaling a limit order book model to millions of parameters, check what information reaches it.
What snapshots leave out
A Level-2 snapshot holds resting volume at ten price levels on each side. The Level-3 stream, called market-by-order (MBO) in the paper, retains the submissions, cancellations and executions that produced it. Once those events are aggregated, snapshots cannot reveal whether a shrinking level was traded through or its orders were pulled. The authors draw on order-flow imbalance work by Cont and coauthors to explain why the distinction could matter for prices.
MBOFormer reads that event history with a 7,203-parameter causal transformer: two blocks, four heads, width 16 and the 128 latest raw events. Each event supplies seven features, including event type, signed side, log tick-depth from the best quote, log size and log elapsed time. MBOFusion adds a small convolutional branch. It samples 16 steps of a 12-dimensional context vector covering drift, volatility, spread, depth, event rates and time to contract close; the full model has 14,371 parameters.
The models predict three-class mid-price direction: up, flat or down. A decision follows every tenth book-changing message, with horizons of 10, 20, 50 and 100 such steps. EPEX continuous intraday supplies 13 delivery days in March 2021, limited to hourly products. The equity samples are 7 days in September 2024 for AAPL and 58 days from August to October 2024 for Bank of America (BAC). The comparison includes DeepLOB, BiNCTABL, MLPLOB with 3.14M parameters on EPEX and 3.07M on the equities, and TLOB at three sizes. One of the two event models leads mean macro-F1 in 11 of 12 market and horizon cells. At h=10, MBOFormer posts 0.651 on EPEX, 0.665 on AAPL and 0.735 on BAC; the 1.20M TLOB posts 0.617, 0.653 and 0.731.
The case for event data is clearest against a parameter-matched model. The 7.2k TLOB receives 16 sampled snapshots spanning about 160 book-changing messages, along with the triggering message and engineered OFI and VWAP features. MBOFormer receives 128 events, or roughly 13 sampled steps, yet beats that TLOB in 11 of 12 settings despite covering less history. A fiftyfold increase in TLOB size, from 23.9k to 1.20M parameters, brings no consistent gain even when its window grows from 16 to 128 rows: at EPEX h=10, the scores are 0.621 versus 0.617. Window lengths differ across models, so representation and lookback are mixed in this comparison. The lookback difference works against the authors' result.
How much of the equity lead holds up?
At h=10, the best L3 model exceeds the best L2 model by 0.009 macro-F1 on AAPL and 0.002 on BAC. Those margins make the 11-of-12 count less persuasive. At BAC h=20, L2 wins: MLPLOB scores 0.592 against 0.589. The abstract calls performance against million-parameter baselines comparable and builds its case around running 5.1 to 8.1 times faster. BAC h=10 seed standard deviations are around 0.001, measuring initialisation noise only. AAPL has a single trading day in its test set; BAC has 14 days. The authors acknowledge that three seeds "do not establish statistical significance across test periods," a consequential limit for equity gaps of 0.009 and 0.002.
EPEX's 0.659 against 0.621 yields a much larger 0.038 gap. Its test covers three days in one month of proprietary data, so public data cannot reproduce the result. The authors suggest liquidity as a possible explanation without claiming to have established it. At h=10, the wall-clock horizon is about half a second for AAPL, five seconds for BAC and two minutes for EPEX, whose median spread is about 60 ticks. The paper cannot identify which market characteristic accounts for the gap.
The label construction also bears on the scores. The flat threshold equals half the mean absolute change, calculated separately for each split and, on EPEX, each delivery contract. Test labels therefore draw on the full test split's distribution. The authors call this "a look-ahead in the label definition, not in any model input," and it applies identically to every model. Comparisons across models retain their footing; absolute performance needs a threshold set using past data, as the paper notes. The models exceed the always-flat benchmark, with 0.651 against about 0.26 on EPEX at h=10.
A further warning sits at the long horizon: MLPLOB falls to 0.389 on EPEX h=100, below the 11.4k-parameter BiNCTABL at 0.454.
The 0.479 ms forward pass
MBOFormer records 0.479 ms median latency on a single-thread M2 CPU, with p99 at 0.791 ms. The 1.20M TLOB takes 2.459 ms and MLPLOB 3.903 ms, making MBOFormer 5.1 and 8.1 times faster, with 113 and 260 times fewer FLOPs. BiNCTABL at 0.187 ms and the small TLOB at 0.301 ms are faster. The authors report both and make a fair case for MBOFormer's quality relative to its latency.
These timings use synthetic inputs. They exclude preprocessing, book reconstruction and decision execution, exclusions the paper states explicitly. Computing tick-depth from the current best quote requires a replayed book at every event. Book maintenance is untimed for both model families, so the measured latency covers model-forward execution alone. The end-to-end difference between maintaining a rolling 128-event window and using one snapshot remains unmeasured.
A desk needs the combined cost of book maintenance and the forward pass.
The paper is direct about returns too: macro-F1 scores "generally do not represent trading revenue." There is no spread crossing or PnL analysis, and a 0.002 BAC gain has no obvious path through a one-tick spread.
One-minute bars cannot replay these events
Our data consist of daily and one-minute OHLCV bars, without order messages, quotes or sub-minute events. On AAPL, a sampled step lasts about 0.045 seconds; one minute contains roughly 1,300 of the paper's decision points. A bar cannot recover the seven event features. The EPEX contracts are unavailable to us as well.
I would want to see the BAC gap across its 14 test days, rather than only the pooled result the paper reports, using a threshold fixed from training data. If 0.009 and 0.002 keep their sign day after day, the classification edge has stronger support; profitability remains a separate question. For now, the useful equity-desk finding is the failed scale-up: fifty times more TLOB parameters, with the window extended from 16 to 128 rows, delivered no consistent gain.