Five standard HFT proxies barely improve the liquidity-supplying regression once the machine-learning measure enters. They add 0.3% to within-stock explanatory power. For liquidity demand, they contribute about 0.6%, taking explanatory power from 0.8% for the model alone to 1.4% jointly. Anyone using quote-to-trade or odd-lot volume as a time-varying HFT control should care about those figures.
Ibikunle, Moews, Muravyev and Rzayev start with NASDAQ's proprietary dataset, which classifies every trade in 120 NASDAQ- and NYSE-listed stocks during 2009 as HH, HN, NH or NN. H identifies a high-frequency firm, while the first letter marks the liquidity demander. Those labels yield two stock-day shares. Liquidity-demanding HFT equals HH plus HN volume over total volume; liquidity-supplying HFT equals HH plus NH over total volume. Their respective means are 0.331 and 0.250, which means labeled HFT participates in roughly half the sample's volume.
Because the labels cover 120 stocks in 2009, the authors use them as training targets. The predictors are 24 daily variables from WRDS TAQ Intraday Indicators, including trade counts, ISO dollar volume, time-weighted depth in shares and dollars, quoted, effective and realized spreads, price impact, lambda, quote-based intraday volatility, retail and institutional order imbalance, and others. Tree ensembles fitted with MSE loss estimate the two NASDAQ shares.
Extra trees lead the comparison. Mean R-squared is 0.805 over ten iterations, versus 0.790 for random forests, 0.783 for a three-hidden-layer network, 0.684 for support vector regression and 0.625 for LASSO. A grid search spanning eight ensemble sizes and eight minimum-split values produces a selected configuration averaging 0.825, with a standard deviation of 0.005. The fitted model is then applied to every TAQ stock-day: 8,314 stocks and 9,440,600 stock-days from 4 January 2010 through 18 October 2023.
The paper builds a measurement instrument. It reports no strategy, returns, costs or turnover. Its practical value to a desk therefore depends on the authors carrying out their stated plan to share the measures.
Why we ran no replication
The training targets come from proprietary NASDAQ trade-level data that attaches an HFT identifier to every transaction. We have no equivalent dataset, and no substitute exists. We also lack the millisecond quote data required by the validation exercises: bid-ask quotes, order submissions or cancellations, message counts and odd-lot trade identification.
A low-frequency proxy built from bars would examine a different object. We therefore ran no substitute test and report no figure of our own anywhere in this article. Every figure below is taken from the paper.
The parts a subscriber can reproduce
Both target variables have explicit arithmetic definitions. Liquidity-demanding HFT is HH plus HN shares divided by total shares for the stock-day. Liquidity-supplying HFT is HH plus NH over that same denominator. A table lists the 24 features alongside their WRDS TAQ variable labels. The authors present that choice as a way to preserve replicability and limit data mining, since a reader with a WRDS subscription can pull identical columns instead of constructing a spread measure by hand.
The estimator is also laid out: extra trees with MSE loss, 10,000 random stock-days per model, ten repetitions, Monte Carlo cross-validation using 75/25 splits, and a grid search across 64 combinations. So is the horse race. The regressions use five conventional proxies, stock and day fixed effects, dependent variables that are standardized, and standard errors double-clustered by stock and day. Training covers January to June 2009; evaluation covers July to December. The resulting sample contains 14,238 stock-day observations for the 120 randomly selected NASDAQ- and NYSE-listed firms with NASDAQ HFT data.
Time separates the samples, while the same 120 stocks remain in each. Model choice and hyperparameter tuning both occur within that single year. The authors state this directly. Much of the introduction and Section 3.3 is devoted to defending extrapolation across the broader market and into later periods.
Judgment calls remain
The first implementation choice appears in the grid-search results. Rank 1 in the optimization table has mean R-squared of 0.814442 and standard deviation of 0.008260, with 5 under split samples and 640 under ensemble size. Yet the text describes a search over 10 to 500 trees and 2 to 50 minimum node splits. Each of the top five rows shows 5 in the first column, while rows 60 through 64 all show 640. We could not reconcile the axes by matching the column names to the stated ranges.
A replicator must make a guess, and the consequence is material. Rank 1 gives 0.814442; rank 64 gives 0.654791. The difference is roughly 0.16 in R-squared.
A second choice concerns venue coverage. The model learns Nasdaq-only HFT shares from consolidated TAQ predictors, then extends those predictions across fourteen years in which Nasdaq's market share changed. Before Table 4, the authors discuss venue aggregation in the form relevant to their event studies. Rerouting by HFTs away from the treated exchange would dilute the measured effect. They consider that concern likely minimal because HFTs favor a stock's primary listing exchange, citing 2023 Amex statistics for time at best prices, quoted depth and spread tightness. Those figures measure market share. The paper supplies no evidence on the separate question of rerouting behaviour. An implementer receives a fixed mapping from a venue-specific label to a market-wide predictor.
The third choice governs how the output should be used. Most of the measure's explanatory power is cross-sectional, as Table 3 shows when the fixed effects change. With day fixed effects alone, the supply measure reaches 74% within-R2, with a coefficient of 0.246 and t of 38.28. The demand measure reaches 50%, with 0.423 and t of 23.23. Adding stock fixed effects cuts those values to 3% and 0.8%.
Conventional proxies deteriorate further. In the liquidity-supplying panels, quote intensity drops from 51% to 0.4%, message count from 53% to 0.7%, and odd-lot volume to 0%. The relative advantage claimed by the authors survives. Still, 3% within-R2 is thin in absolute terms for a variable intended to condition a time series. Their own tables establish that limit.
Speed bumps, feeds and stale quotes
The two construct-validation exercises deserve close inspection before anyone trusts the series. Within-R2 ranges from 1.3% to 11% across those difference-in-differences tables, and both studies use 10-working-day windows.
Amex introduced a 700-microsecond round-trip speed bump on 24 July 2017. Relative to pre-event means, demanding HFT falls 2.8%, with coefficient -0.005 and t of -2.34. Supplying HFT declines 4.6%, with -0.007 and t of -3.31. N is 45,530, and within-R2 is 1.3% and 3.5%. In October 2011, Nasdaq reduced data dissemination latency from 3ms to 1ms. The two measures then rise 0.7% and 1.1%. Both coefficients are 0.002, with t of 2.12 and 2.10, and N of 43,234.
Those movements have the predicted direction: the measures increase after the upgrade and decrease after the speed bump. Effect sizes range from 0.7% to 4.6%. The authors regard the difference between events as informative. A speed bump directly increases latency, whereas the feed upgrade changes only the consolidated tape, which HFTs can bypass by purchasing direct feeds.
The larger decline among suppliers is the more interesting result. The authors interpret it as complementary to Aït-Sahalia and Sağlam, who report wider quoted spreads and lower liquidity after the same speed bump.
The latency-arbitrage exercise produces strong t-statistics on a narrow cross-section. Stale-quote opportunities average 68 per stock-day, with a standard deviation of 169. A one-standard-deviation increase accompanies a 1% rise in demanding HFT, coefficient 0.018 and t of 3.78, alongside a 1.6% decline in supplying HFT, coefficient -0.020 and t of -2.02. Snipers enter; market makers retreat. The cited theory from Budish, Cramton and Shim; Foucault, Kozhan and Tham; Aquilina, Budish and O'Neill predicts exactly that pattern, since arbitrage opportunities promote liquidity demand and deter liquidity supply.
These regressions cover 246,139 stock-days. Their universe contains only 120 randomly selected NASDAQ- and NYSE-listed firms, the same names used for the training labels, because processing the quote data over the full sample was computationally prohibitive. The measures therefore move with theory on those 120 names. The exercise says nothing directly about the remaining 8,194 stocks in the applied sample.
A nearby result is genuinely surprising. The authors add millisecond quote variables such as message counts, quote update frequency, sub-100-share volume and 100ms midpoint variation. Repeating the July-December 2009 out-of-sample comparison moves R-squared only from 82% to 84%. The resulting measures correlate 0.99 and 0.96 with versions built from daily indicators.
Correlations among the authors' own predictors provide their explanation. Total trades and message count correlate at 0.90. Message count exceeds 0.65 with both ISO trades and market depth; quote revision frequency exceeds 0.70 with trade frequency, ISO trades and depth. Anyone paying for message-level feeds solely to construct an HFT control should read that result twice.
Earnings produce opposite signs
The application uses Weller's jump ratio, defined as cumulative abnormal return over trading days [-1, 1] divided by cumulative abnormal return over [-21, 1]. A larger ratio indicates that less information arrived beforehand. Raising demanding HFT from its 25th percentile, 0.222, to its 75th, 0.414, increases the ratio by 6.6% of its mean, with coefficient 0.178 and t of 4.57. Raising supplying HFT from 0.131 to 0.259 reduces it by 3.3%, with -0.133 and t of -2.71.
The sample comprises 49,515 firm-quarters for all U.S.-listed common stocks from 2010 to 2023. Future earnings response coefficient tests point the same way in the specification with controls: the supply interaction is 2.676, with t of 5.25, and the demand interaction is -2.018, with t of 4.56. N is 157,343.
Within-R2 reaches 0.4% for the jump ratio and 4% for the FERC. The authors explicitly say that causal identification here is challenging. Their caution is warranted.
The reconciliation gives the finding more weight than a two-coefficient result would carry alone. Weller's MIDAS proxies load in opposite directions on the two measures. Using all U.S.-listed common stocks from 2012 to 2023, with N of 43,091, the odd-lot ratio loads at +2.714 on demand and -2.343 on supply. Cancel-to-trade gives +0.839 and -1.133; trade-to-order gives -1.208 and +1.340. Every t-statistic exceeds 10 in absolute value.
The authors interpret these estimates as evidence that Weller's measures mainly capture liquidity-demanding HFT in his sample. Under that reading, his negative finding and their positive one reflect the same underlying fact through a coarser measure. Anyone with MIDAS can check it.
The strongest counterweight also comes from the paper. The authors repeat the jump-ratio and FERC exercises using the raw 2009 NASDAQ labels for the 120 randomly selected stocks with NASDAQ HFT data. Neither produces a result. Jump-ratio coefficients are 0.997, with t of 0.52, and -0.903, with t of -0.56. N is 466 firm-quarters, falling to 401 in the FERC version. Limited sample size is their explanation, and 466 observations make it plausible. Yet the only setting where HFT is observed directly does not corroborate the headline. The finding appears only in the extrapolated series.
A stock holdout would answer more
The existing data allow one especially useful test: train on 90 of the 120 names, predict the remaining 30, then report within-stock R2 for stocks unseen by the trees. Every holdout in the paper is temporal. One divides January-June from July-December 2009. Another uses a near-training period from January 2010 through December 2012. During that later window, the jump-ratio estimates remain, at 0.114 with t of 2.59 and -0.101 with t of -2.10. The FERC results also persist, at -3.982 with t of -3.41 and 3.666 with t of 2.81.
None withholds an individual stock. Yet the applied cross-section is 69 times larger than the training cross-section.
Until that experiment is run, the tables support using these measures as a cross-sectional HFT ranking. The supply measure explains 74% of within-day variation, while conventional proxies peak at 53%. The time-series component deserves directional use rather than calibration.
We previously objected to a radius selected in hindsight and a universe chosen over the full sample in our note on Wasserstein portfolios. This case has the same general shape in milder form: model selection and hyperparameters come from the sole labeled year available. The authors acknowledge the extrapolation assumption in the introduction and support it with prior work using the same NASDAQ dataset. A stock-level holdout among the 120 names would resolve more than another citation.