AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

No model cleared the 21-bps fee floor

Jadouli's frozen crypto selector lost 6.72% across 19 July cycles; the extrema arithmetic lasts.

2026-09-08 · 8 min read

Reviewing: Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models · Ayoub Jadouli · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Cost-Aware Selective Extrema Crypto Timing

Backtest period 2020-01-01 to 2025-10-08 · hypothetical, net of modelled costs

Sharpe
0.00
Total Return
0.0%
Max Drawdown
0.0%
CAGR
0.0%
Volatility
0.0%
Trades
0

What the paper reports for its own strategy

  • Frozen mandatory-daily selector: -6.72% compounded, 2026-07-01 to 07-19, 19 cycles, 3 wins/16 losses, 31 bps completed-cycle assumed cost, mean -36.40 bps/cycle
  • Frozen mandatory-daily selector cost stress, same period: -4.74% at 20 bps, -6.72% at 31 bps, -10.21% at 51 bps (terminal 95.2579 / 93.2815 / 89.7873 USDT from 100)
  • Local-minimum selected model: -1.79% compounded, 2026-07-01 to 07-12, 9 cycles (5/4), 31 bps, max drawdown 2.07%
  • Local-maximum selected model: -2.80% cash-cycle advantage versus continuous holding, 2026-07-01 to 07-12, 15 cycles (8/7), 31 bps, max drawdown 2.91%
  • Daily paired-extrema adaptation: -44.30% versus -41.20% cost-matched buy-and-hold, 2025-07 to 2026-06, 7 cycles (1 win/6 losses), 31 bps round trip applied as 15.5 bps per side; -44.10% at 21 bps
  • Slow rotation baseline: +313.01% compounded, 2022-10 to 2026-06, 45 scored months (21 traded, 15 target-active), 31 bps round trip charged as half-cost per unit one-way turnover, max drawdown -48.38%; equal-weight hold +78.69%

Across 28 fitted models in two directions, every gross mean cycle advantage fell short of the 21-bps fee floor. The high-water mark was +20.0 bps.

Everything else in the paper explains why that shortfall kills the trade.

Jadouli presents the paper as a negative self-audit, down to its title: predictive extrema, unprofitable policies. Five campaigns tested candle-based Binance Spot timing models through deterministic simulators, using scripted fixed-seed runs and paper trading. Every campaign reached the same operational verdict. A classifier may rank rare future extrema well while the stateful policy built from those rankings still loses money after costs.

Rare-event timing was meant to supply the economics. A five-minute bar receives a local-minimum label when its low is the extremum in a centered 49-bar window and the following 24 bars produce at least a 40-bps rebound. The classifier uses 48 bars of causal OHLCV, activity, indicator, price-location and clock features. It buys at the next open with a +60/-50 gross bps bracket. The maximum direction mirrors that setup: sell an already-held Spot position into cash, then repurchase later. Its reported return measures advantage versus continuous holding rather than short P&L.

All data are Binance Spot klines. The five-minute set covers Three pairs, BTC, ETH and SOL against USDT, from March 2025 to July 12, 2026. Models fit through March 2026, validate from April to June and evaluate from July 1 to 12. The one-minute set contains Ten USDT pairs from March 2025 to July 19, 2026. Daily candles for BTCUSDT and ETHUSDT run from June 2021 through the July 1, 2026 terminal open. Minima have 1.418% prevalence; maxima have 1.603%. Every evaluated policy finishes at NO_TRADE.

The paper also draws a clear boundary around its AI-assisted audit. Agents performed literature retrieval, separately tasked critique passes, artifact reconciliation, documentation and source packaging. They made no trading decisions. Jadouli says the process caught an overstated holdout label and silent plot fallbacks in an earlier draft, setting up the later forensic examination.

A second construction forces a daily choice among ten USDT pairs using one-minute bars. One ExtraTrees realized-net regressor ranks the ten-pair universe after the candle timestamped 16:00 UTC has completed. At the 16:01 open, the policy must enter exactly one pair. It remains open for at most four hours, with a +100-bps gross target and a -150-bps gross stop. The Third campaign adapts Gurgul, Lessmann and Härdle's paired-extrema idea to daily BTCUSDT and ETHUSDT using OHLCV alone, excluding their blockchain, GitHub, search, macro and text variables. The Fourth is a consumed monthly rotation control spanning 2021-06 to 2026-07. The Fifth revisits an archival shared ten-pair LSTM that the author calls One4All.

Costs come from assumptions, rather than measured commissions. The primary 31-bps label enters the three setups differently. Intraday campaigns deduct it as a flat amount per completed cycle. The paired daily study charges 15.5 bps multiplicatively on each executed side, including terminal liquidation. The rotation control applies half-cost per unit of one-way turnover. Stress levels vary too: 20/51 bps for the mandatory selector, and 21/51 bps for local extrema and paired daily. Binance's published regular-user Spot schedule gives 0.100% per maker or taker side. The 21-bps floor combines two 10-bps sides with a 1-bps buffer.

Jadouli concedes the central arithmetic in the abstract. Gross mean advantages are reported at 11.11 and 12.21 bps, below even the 21-bps stress. The abstract concludes that event-ranking performance did not establish positive executable policy value. The title follows its reasoning. Because this is a negative study by design, the useful question is what survives after the author has already surrendered the headline result.

The edge disappears in execution

The selected local-minimum model, an attention CNN-LSTM, recorded ROC AUC 0.729 and average precision 0.043. Its held-out set had 10,293 rows and 146 positives. Across 9 cycles, it returned -1.79%, with gross mean +11.1 bps and net mean -19.9. On the maximum side, AUC was 0.750 and AP was 0.053. The policy returned -2.80% versus continuous holding over 15 cycles, with gross mean +12.2 bps and net mean -18.8. Costs of 10.9616 and 12.0690 bps would have set those realized paths to zero. The paper properly treats them as realized-sample break-even summaries from 9 and 15 cycles, rather than forward estimates.

The surrounding panel makes the result structural. During the July 1 to 12 model-specific evaluation, the strongest gross mean among all 14 minimum-direction models was +20.0 bps over 18 cycles. It came from a price-location attention hybrid with AP 0.042. Across all 14 maximum-direction models, the selected attention CNN-LSTM led at +12.2 bps over 15 cycles. A GRU was the best unselected model, at +8.1 bps over 41 cycles. Both remain below two 10-bps sides plus a buffer.

Rank quality also moved against trading performance. The logistic detector reached AUC 0.973 and AP 0.360, then returned -4.02% over 15 cycles. The maximum-side histogram gradient boosting (HGB) model reached AUC 0.969 and AP 0.288 while losing 16.50% over 45. The paper marks both rows as post-hoc diagnostics, and neither passed the validation gate. Their message still matters: better annotation of completed extrema generated more trades with poorer per-trade economics.

The daily adaptation reaches the same conclusion with more label observations. Across 704 label-valid held-out symbol-days, minimum AUC/AP/prevalence was 0.8742/0.1341/2.983%, while maximum AUC/AP/prevalence was 0.8962/0.1158/2.131%. The portfolio lost 44.30% over seven cycles, compared with -41.20% for cost-matched equal-weight BTC/ETH. It produced one win and six losses, ending at 55.70 USDT versus 58.80 from 100. Reducing assumed cost to 21 bps shifted the strategy only to -44.10%. Fees were not the binding constraint in that experiment.

Can 19 cycles support the abstract?

The paper describes the frozen ten-pair selector as its strongest later-period evidence, conditional on extensive predecessor search. At 31 bps, it lost -6.72% over 19 July 2026 cycles, with 3 wins and 16 losses. Mean performance was -36.40 bps per cycle. For the disjoint July 8 to 19 extension, the same stored model was reloaded and identified by hash, without retuning. It lost 354.28 bps on 3 wins and 9 losses. Every cost path ended negative: 95.2579, 93.2815 and 89.7873 USDT from 100 at 20, 31 and 51 bps.

I would update little from those observations. The descriptive Clopper-Pearson 95% win-rate interval runs from 3.4% to 39.6%. Neither the reporting cutoff nor the sample size was preregistered, as the paper acknowledges. Concentrated selections explain much of the path. In a descriptive attribution, TRXUSDT appeared in 9 of 19 forced cycles, won once and totaled -334.28 net bps. ETH's two cycles totaled +138. The freeze itself was clean. For the extension, no universe, feature, target, stop, time, threshold, model, training cutoff or cost changed. Still, its 19 observations remain 19 observations. The paper also states that the frozen summary benchmarks only against cash and no-trade.

At least 946 candidates were enumerated across eight campaign summaries and compared from April to June before the freeze. April to June was later reused in the final fit. The author's phrasing is therefore prospective conditional on that history. Each extrema campaign also consumed 238 validation policy comparisons per direction. Beyond those sit 240 rotation policies and 3,996 profiles in the archival sweep. A win-rate interval spanning 3.4% to 39.6% leaves a single 19-cycle result at the end of that funnel unable to distinguish a dead strategy from a noisy sample. Interpretation of those 19 days cannot settle it.

One positive path, heavily concentrated

The consumed rotation control compounded +313.01% across 45 scored months at 31 bps, compared with +78.69% for equal-weight hold. Targets were active in only 15 months, and 8 of 21 traded months were positive. November 2024 returned +160.422%, accounting for 67.48% of net log gain. Remove that month and terminal wealth drops from 413.01 to 158.59 USDT.

A predeclared two-asset cap reduced maximum drawdown from -48.38% to -38.69%, while lowering wealth to a terminal 350.22 USDT. In the paired comparison, cap-minus-baseline log return is -0.164925. Its one-sided 90% moving-block lower bound is -0.738153, leaving the cap's cost unresolved. Absolute-return selection elsewhere shows the same pathology on a smaller scale. The selected 14-day HGB profile compounded +73.93% over seven folds, yet its median fold excess versus buy-and-hold was -18.60 percentage points.

The forensic section deserves imitation. Jadouli downgrades his own 30-day holdout for the archival LSTM to a repartitioned archival diagnostic. Those dates had shaped earlier architecture work, the four-hour label horizon remained unpurged at split boundaries, the decision and entry fill used the same close, and raw result directories are missing. An earlier publication script had also inserted hard-coded summaries and synthetic equity paths. Those were removed. Downgrading a headline result while naming the mechanism carries more weight than most positive findings at this sample size.

What abstention would need to clear

Jadouli sets the limits of his claim. Every conclusion is specific to the exchange, universe, period and simulator. Section 8.3 confines the result to the contextual statement that these intraday multi-pair extrema and mandatory-action policies did not establish positive value. The paper cites Bysik and Ślepaczuk for comparison. They test hourly BTC-USDT with XGBoost, LSTM and iTransformer across 27 walk-forward folds. Their naive sign policies fail at 10 bps, while a cost-aware magnitude filter restores profitability in selected configurations. That boundary is fair.

Inside this universe, it changes nothing. Break-even costs of 10.9616 and 12.0690 bps remain below a two-sided 10-bps fee floor. That arithmetic is the transferable result.

The mandatory daily selector excludes abstention by construction, as the paper states. The extrema policies could abstain. Each direction consumed 238 validation policy comparisons across up to 14 threshold quantiles, yet no window or model cleared the predeclared gate for positive return, mean return and drawdown. Jadouli argues that further threshold search on those consumed periods would not answer the question. We reached a related observation from the opposite direction in our note on Wysocki's 0DTE ranker: the abstention layer never bound out of time, so it added nothing.

Arithmetic sets the design bar before statistics enter. A gross mean per cycle must clear the 21-bps floor and leave room above the 31-bps primary assumption. The test needs a point-in-time universe, with label horizons purged at both split boundaries. Only then does the cost assumption become an open question. Until that evidence exists, this paper remains a well-documented account of an edge too small to pay Binance. The frozen 19 cycles contribute the least.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each walk-forward month:
  1. Train on prior data after purging/embargoing label horizons.
     - Classifier A: probability current bar is near a local minimum.
     - Classifier B: probability current bar is near a local maximum / adverse exit risk.
     - Regressor: next executable open to 4h-later open return in bps.
  2. Calibrate classifier probabilities on the calibration fold only.
  3. Select entry thresholds on calibration data subject to:
     - positive calibration net return after costs
     - positive mean trade expectancy
     - max drawdown below required gate
     - minimum calibration trade count
  4. If the policy gate fails, do not trade that test fold.

On each completed 5-minute bar during the test fold:
  Build OHLCV-only features using past data only.
  Predict:
    p_minimum, p_maximum_risk, expected_4h_return_bps
  Compute predicted_net_edge = expected_4h_return_bps - 31 bps primary round-trip cost.

Entry rule when flat / capacity available:
  Enter long at next resampled bar open if:
    p_minimum >= selected minimum-probability threshold
    predicted_net_edge >= selected net-edge threshold
    p_maximum_risk < 0.35
    next execution bar exists
  If multiple symbols qualify, choose highest edge minus cost.
  Size as equal-weight capped position, max 10% NAV per symbol, max 3 concurrent positions.

Exit rule for each open position:
  Exit when any condition is true:
    stop loss hit: -75 bps
    profit target hit: +120 bps
    timeout reached: 48 primary 5-minute bars
    p_maximum_risk >= 0.55
    expected net edge <= 0 bps
    symbol/execution price unavailable
  Evaluate stop/target using resampled bar high/low; if same-bar tie, assume adverse stop first.
  Apply half round-trip cost per side, including terminal liquidation.