The headline 0.82 AUC mostly looks like volatility clustering against a fixed threshold. The lasting contribution is the taxonomy: four formal definitions of an intraday downside event on one-minute bars, three of them new. The paper's own Jaccard matrix confirms that they identify different minutes. Kończal and Połoczański acknowledge the strongest objection in Section 5, directly in the text rather than burying it in a footnote.
We could not run any part of the study on our own data. This review contains neither a replication nor a test of the paper's claim, and no figure of our own appears anywhere in it. We hold no Xetra or Nasdaq Stockholm crypto ETP instruments, so we cannot reproduce the traded universe or the reported AUC values. We also lack bid/ask quotes, effective spreads, bid/ask volumes, dark-pool volume, periodic-auction volume and order-flow-imbalance fields. All 12 microstructure features are absent. We would have access only to OHLCV-derived volatility, drawdown, return and volume features.
Venue-specific ETP prices for the same product are unavailable as well. Cross-venue divergence is therefore beyond what we could examine, leaving only the temporal price-path mechanisms, no-recovery and momentum-reversal. Our price history ends in 2024, whereas the paper's sample extends into December 2025.
Kończal and Połoczański study European crypto exchange-traded products (ETPs). These instruments trade on regulated venues during exchange hours, while their underlying coins trade around the clock. The starting universe contains 200 ETPs listed across Europe. Xetra is the primary venue for 46.6%, followed by Nasdaq Stockholm at 17.3%, SIX Swiss at 16.8% and Euronext Paris at 13.6%.
Data completeness reduces that universe to four instruments: VanEck Bitcoin and Ethereum on Xetra, plus Virtune Bitcoin and Ethereum on Nasdaq Stockholm. The sample spans January 2024 to December 2025, using one-minute bars from the continuous session and dropping zero-activity bars. Final sample sizes range from 155,510 bars for VBTC.XE to 73,448 for VIRSETHS.ST. Minute-return standard deviation climbs monotonically across the four, from 0.1229% to 0.3698%. Excess kurtosis lies between 27.55 and 58.48, while the worst single minute is -14.133%.
Where would the money come from? The paper ends with classification. It presents no strategy, P&L or cost calculation, and claims none. We see the economic channel in execution and risk: a minute-ahead warning that a fall is unlikely to reverse could justify widening or delaying a child order. This is our extrapolation. The paper does not test it.
Its microstructure results are consistent with that channel. At DPOT bars, the median effective spread ratio relative to normal bars reaches 3.99 and Kyle's lambda reaches 11.24 for VIRSETHS.ST. The corresponding VBTC.XE figures are 2.05 and 3.42. Relative volume peaks around no-recovery bars at 2.67 to 5.08 times normal levels. For three of the four definitions, median order imbalance changes from positive to negative at anomaly bars across all four instruments.
Four labels, different minutes
- DPOT, the benchmark, marks a return below a quantile obtained from a generalised Pareto fit to left-tail losses. The threshold is the 95th percentile of training losses, with target rate p = 0.01. The authors flag that the quantile is estimated on the training sample and then applied across the whole series.
- DCV, cross-venue divergence, uses the absolute residual from a rolling no-intercept OLS of Stockholm returns on Xetra returns. Its window covers 1440 active bars, and the cutoff is the rolling 99th percentile. Both members of a pair receive the same labels.
- DNR, no recovery, requires a return below the rolling 1st percentile. Cumulative return over the next K = 10 active bars must also remain below 0.3 times the absolute initial return.
- DMR, momentum reversal, combines a return below the rolling 1st percentile with positive cumulative return over the previous 5 bars.
Every definition fires on under 1% of valid bars. DPOT ranges from 0.380% to 0.599%, while DNR ranges from 0.791% to 0.929%. DCV's Jaccard overlap with the others is only 0.012 to 0.030. DNR against DMR reaches 0.319 to 0.467. Two definitions come close to restating each other; one captures a separate object.
Why the fixed line wins
Across 80 instrument-definition-classifier-scheme combinations, AUC runs from 0.558 to 0.823. DPOT occupies the top end. Logistic regression records 0.821 for VBTC.XE with cumulative features and 0.823 for VIRBTC.ST. DNR produces 0.577 to 0.732, while DMR produces 0.558 to 0.764. The composite label, triggered by any of the four, returns 0.633 to 0.740.
DPOT takes its threshold from the training sample and holds it fixed. DNR and DMR instead use a rolling 1440-bar percentile, allowing their thresholds to move with volatility. Recent volatility can therefore predict a breach of DPOT's fixed level much like a conditional scale forecast.
Feature importance supports that reading. roll_std_15 and drawdown_15 appear among the top five for three of four instruments under DPOT. Plain price and volatility features outperform microstructure features throughout. Logistic regression also matches or beats XGBoost and LightGBM for most combinations. The like-for-like prediction-table comparison is especially clear: for VBTC.XE with cumulative features and logistic regression, DPOT scores 0.821, DMR 0.762 and DNR 0.732. The gaps are 0.059 and 0.089.
The operating point imposes a harder limit. F1-optimal thresholds span 0.436 to 0.950, and precision "rarely exceeds 0.10" across all models and definitions. The authors give the problem its full weight: "since the thresholds were chosen using labels from the test set, these results are only descriptive and should not be interpreted as genuine out-of-sample performance."
Credit for putting that in print. They return instead to the separation metric: "These results justify the use of AUC, ROC as the main evaluation metric. AUC, ROC evaluates how well a model separates the two classes across all classification thresholds, while precision and recall depend on the threshold selected." AUC works as a ranking diagnostic, exactly as the quotation argues. Position sizing still needs a validated threshold. Readers who retain the headline that all four anomaly types are "predictable one bar ahead, with AUC-ROC values of up to 0.82" may easily forget the precision sentence.
What is visible at bar t?
The paper explicitly lags the features by one bar and keeps them causal. DMR's momentum gate requires positive cumulative return over the five bars preceding the event bar. In our reading, the classifier can already see much of that gate through its lagged inputs. The remaining forecast is mainly the extreme-return leg. Under DMR, rsi_5, vol_ratio and roll_mean_5 rank in the top five for all four instruments. A model partly learning the deterministic gate should produce exactly this pattern.
DNR faces the opposite timing issue. Its forward ten-bar recovery condition remains unknowable for ten active bars, as the authors state plainly. Within the paper's AUC ranges by definition, DNR has the lowest ceiling among the four single-anomaly targets at 0.732. The others are DPOT at 0.823, DCV at 0.774 and DMR at 0.764. A single separation metric says little about a label whose outcome is partly unresolved when the prediction is made.
We have encountered this family of problem before, in an index whose attribution layer collapsed to a row sum (/articles/the-triadic-stress-index-ties-the-measure-it-set-out-to-beat).
Cross-venue divergence at one-minute frequency
The paper reports one-minute cross-instrument correlations between 0.174 and 0.435, compared with 0.745 to 0.982 daily, and identifies the Epps effect. We read the low-frequency match differently from the one-minute result. At one-minute frequency, the series barely co-move. A rolling no-intercept regression can therefore leave asynchronous printing as a large part of the residual, and the 99th percentile of its absolute value may capture staleness as readily as dislocation.
Two reported patterns support this interpretation. DCV has a weak order-imbalance signature whose sign varies, with ratios of 0.25, 0.47, 0.12 and -0.04. The other three definitions change sign consistently. Amihud and Kyle lambda ratios at DCV bars fall below DNR and DMR on Xetra, then rise above them on Stockholm, reversing the venue effect.
DCV bars remain illiquid. Effective spread ratios range from 1.21 to 1.81, while Kyle lambda reaches 7.15 and 6.30 on the two Stockholm listings. Given the sign-inconsistent imbalance and the Epps-effect correlations, I read most DCV events as stale prints rather than tradable dislocations. Quote midpoints on a synchronised clock would change my view if the residual tail survived.
Four instruments out of 200
The authors select instruments by the largest count of non-zero one-minute bars within the 200-ETP universe over the full period. The resulting cross-section is conditioned on liquidity and survival. Effective independence is also below four because paired listings share an identical DCV label.
Evaluation relies on one chronological split. Classifiers are "trained once on the first 80% of the data and tested on the remaining 20%, without retraining." A single hyperparameter set applies throughout because "the same hyperparameters are used for all instruments and prediction targets rather than being tuned separately for each model. This reduces the risk of overfitting." The choice does reduce that risk. It leaves no walk-forward evidence that the models are well specified.
Two years cover one directional cycle. The cumulative feature scheme wins 68 of 80 comparisons against the session-based version, by roughly 0.009 AUC on average. I would not commit a design choice to a gap that small.
Someone with the necessary data could test the no-recovery and momentum-reversal mechanisms on spot minute bars. That would be a different study from this paper.
A volatility-adjusted threshold for DPOT, paired with an operating point selected on the training set rather than the test set, would make me revise this view. If DPOT's 0.82 remains, the paper has found something beyond conditional scale.