AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Sheikh and Bhat on crowding and anomaly half-lives

A theoretical saturation result, a simulated detector, and our SPY adaptation

2026-10-05 · 6 min read · Regime-conditioned momentum · US ETFs

Reviewing: Pockets of predictability: a stochastic crowding model and finite-sample regime detection for the endogenous decay of trading anomalies · Sajad Ahmad Sheikh and Dilawar Ahmad Bhat · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

SPY Positive Momentum with a Walk-Forward Predictive-Covariance Gate

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

Run on a different market than the paper

The paper trades cta; we have no price data for it, so the strategy below runs on US ETFs instead. The judge did not state why the mechanism survives the swap.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (3)

  • eqs 25–28, predictive regression and detector: R_{t+1} = β₀x_t + βs_tx_t + ε_{t+1}; r_{t+1} = R_{t+1} − β₀x_t = βs_tx_t + ε_{t+1}; p_i = x_{i−1}r_i = βs_{i−1}x_{i−1}² + x_{i−1}ε_i; T_t = (1/n)Σ_{i=t−n+1}^t p_i; ŝ_t = 1{T_t > τ}, τ = β/2.: Use completed raw-SPY-momentum/next-day-return pairs, their centered sample covariance, and a frozen training-median threshold; trade binary long-or-cash weights. (invalidates: The paper's Gaussian-product detector interpretation for the SPY gate; its known-β midpoint rule; direct application of its detector-error and performance-transfer guarantees to this backtest)
  • The paper removes the possibly nonzero baseline β₀x_t before computing its products. Its statistic is an average of products for an idealized unit-variance, zero-mean Gaussian predictor, not automatically the centered sample covariance of raw SPY momentum and returns.: Compute centered covariance of raw SPY momentum and subsequent SPY returns, without claiming it is the baseline-adjusted Gaussian-product statistic. (invalidates: Identifying SPY's covariance estimate with T_t; applying the Gaussian-product detector-error bound to the SPY gate)
  • The paper’s theoretical strategy uses signed, continuous weight x_t; the requested strategy must be long SPY only when momentum is positive and the covariance gate is open, and must otherwise hold cash.: Use binary SPY-or-cash weights, with no short positions. (invalidates: The paper's continuous-weight oracle-versus-unconditional Sharpe formulas as quantitative predictions for either SPY strategy; the paper's continuous-weight turnover implications)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 13.5%Jul 2024
Sharpe
0.33
Total Return
13.5%
Max Drawdown
-10.2%
CAGR
2.9%
Volatility
9.1%
Beta vs SPY
0.17
Trades
90

What the paper reports for its own strategy

  • Detector-conditioned strategy (Monte Carlo, synthetic data only): per-period Sharpe 0.224 at β=0.4, window n=100, q̄=10^-3, π1=1/2, σ_ε=1. This captures 72% of the oracle-minus-unconditional gain. No transaction costs are modeled, and there is no real sample period.
  • Same strategy, window sweep (synthetic): peak per-period Sharpe 0.217 near n=100 at β=0.4. No costs; no real sample period.

Once an anomaly is crowded enough, more competing capital should leave its decay rate unchanged. That is the trading claim in Sheikh and Bhat. The regime detector built around their model captures 72% of an oracle's simulated Sharpe gain, though the simulation was designed to contain predictability. The paper makes no empirical claim: every result comes from closed forms or Monte Carlo on the authors' own model. They put the limit plainly: "the decisive test of the framework is empirical."

Why decay stops speeding up

Capital deployed against an anomaly reduces its alpha linearly. It moves at speed κ toward a level justified by a lagged alpha estimate. The estimate is an exponential smoother updated at rate θ; 1/θ represents the look-back in allocators' backtests and track records. Flows unrelated to the anomaly's merit supply the noise. These pieces form a two-dimensional linear diffusion with a closed-form solution. Its stationary alpha is Gaussian, centred on c, the marginal cost of exploitation. The authors describe this as a dynamic Grossman-Stiglitz equilibrium.

The dividing line is θ = 4λκ, with λ representing price impact in reduced form. Above it, alpha returns monotonically to cost. Below it, capital arrives on stale estimates, overshoots, withdraws and comes back. Alpha then cycles around cost without an exogenous regime process. At the baseline calibration of λ=0.5, κ=2, θ=1, σ=0.06 and c=0.02, in annual units, the cycle takes about 7.3 years. The envelope half-life is 2log2/θ, about 1.39 years.

That half-life carries the paper's sharpest prediction. Within the oscillatory regime, θ alone determines it: "faster or larger competing capital does not accelerate the decay of the anomaly". Decay across crowded anomalies should track evaluation horizons while staying flat as more capital pursues them. The paper does not run this empirical test. It identifies itself as entirely theoretical and simulation-based; the open anomaly datasets it cites could be used to test the prediction.

The linear model also assigns negative alpha a probability of 0.2525 at baseline, a problem the authors address with multiplicative depletion. Over 20,000 simulated years, that version keeps alpha positive. Its mean is 0.01978, close to c, while its standard deviation reaches 0.03758 and its median is 0.00670. The Gaussian law disappears. A mean centred on cost survives globally; the dichotomy and saturation survive in the linearization around cost, after matching marginal impact there (λ_m·c = λ). The detector uses the linear model, which is the model followed below.

When is the pocket open?

The paper turns the diffusion into a binary state by sampling alpha and calling the pocket open above c+δ. With weekly sampling and δ=0.01, the open fraction is 0.36944. Closed form gives a one-step switching probability of 0.04198, against 0.04207 in simulation.

The failed approximation is more revealing. Match a Markov chain to the exit probability and its mean open run comes close: 17.60 weeks versus 17.49. Across the first 120 weeks, however, its survival curve misses the thresholded process by as much as 0.352. That gap matters for anyone sizing a regime filter with a two-state HMM.

The detector starts from R = β0·x + β·s·x + noise, where s marks the regime. After removing the baseline loading β0, it averages x times next return over the trailing n periods. An average above β/2 declares the pocket open; β/2 lies midway between the high-state mean β and the low-state mean of zero. Theorem 3 bounds misclassification using the chance of a switch within the window and two exponential noise terms. Every constant is explicit.

An ETF momentum application requires x, the β0 to remove, the β setting the threshold and σε for the window rule to be fixed in advance. The main bound assumes β is known. An estimated-β version adds a confidence-failure term and reduces the separation from β to β minus an error radius r that holds with stated confidence. We did not find a simulation of that version.

The window problem

Minimising the bound yields n = γ⁻¹log(4γ/q̄). Here q̄ is the switching probability, and γ is the slower noise rate. A short window admits noise; a long one spans more crossings. The authors acknowledge how loose the bound can be. With q̄=10⁻⁴ and β=0.5, empirical error reaches its minimum near 1% at about n=200. The prescription is n ≈ 1.3×10³, carrying a guaranteed error of 0.16. As the paper says, n* "should be interpreted as a conservative certificate-driven default, not as a prediction of the empirically exact minimizing window."

The abstract strains that distinction. It says the detector recovers roughly seventy percent of the oracle gain "at the theoretically prescribed window length". The Sharpe experiment instead uses n=100, chosen from simulations. For q̄=10⁻⁴ and β=0.5, the formula gives 1.3×10³.

When regimes shorten, accuracy falls too. At q̄=10⁻³, the lowest error is about 6% for β=0.5, 10% for β=0.3 and 16% for β=0.2.

A simulated Sharpe gain

For β=0.4, q̄=10⁻³ and π1=1/2, the unconditional strategy earns a simulated per-period Sharpe of 0.184 (theory 0.183). The oracle reaches 0.240 (theory 0.239); the detector reaches 0.224 at n=100. The oracle observes the true regime, which investors cannot. It suppresses dormant-state trades that contribute variance (1-π1)σε² without adding mean. On these figures, the detector captures 72% of the oracle gap.

The number needs qualifications. At the same β=0.4, the figure's window sweep shows a peak of 0.217 near n=100, with an unconditional 0.175 and an oracle 0.234. The small inconsistency deserves reconciliation. Theorem 4's lower bound for these parameters is printed as "≥0" in the summary table, offering no guarantee here. The detection and Sharpe experiments use a symmetric Markov chain that the paper calls a transparent benchmark. The theorems concern the threshold process, yet no Sharpe experiment uses it. In the weekly bridge experiment, a geometric-sojourn chain misses that process's survival curve by up to 0.352. These Sharpe ratios exclude costs.

Our SPY construction

The paper places no trades; its detector operates on simulated regressions. We built a separate construction from it using one US ETF, SPY, from 2020-01-14 to 2024-07-01. Our single automated run tests our implementation and says nothing about the authors' work. It holds SPY when 20-session momentum is positive and centred covariance between past momentum and next-day return clears a threshold. The threshold is the median of estimates from the previous 252 sessions, held fixed for each 21-session block. Otherwise, the position is cash. We specified three covariance windows (21, 63 and 126); the run does not identify which one generated the figures below.

After a $0.004-per-share commission, our run returned 13.45% in total. Its Sharpe was 0.33, volatility 9.14% and maximum drawdown -10.15%, across 90 trades. Beta to SPY was 0.17, consistent with a part-time long SPY position. We did not run a date-matched ungated momentum baseline, leaving the gate's contribution unestablished. We specified next-open fills, while the run used market-on-close prices. The discrepancy could move the results either way.

The simulated detector receives a predictive loading set by design at β=0.4 against unit noise. Daily SPY momentum offers nothing close to that signal-to-noise. Our construction differs from the model as well. Its rolling-median threshold excludes about half the windows regardless of regime, so we would expect substantial misclassification. Its long-or-cash position also gives up the short side and magnitude weighting responsible for the model's mean profit.

Regime-conditioned versions of standard anomalies would make the detector worth trading if they earn higher Sharpe ratios out of sample, as the paper predicts. The saturation claim has a separate empirical test: post-publication half-lives in the open anomaly datasets cited by the authors should cluster around institutional evaluation horizons and remain flat as the capital chasing them grows.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each covariance window n in {21, 63, 126}, separately:
  Divide eligible sessions into sequential 21-session test blocks.
  Before each block, calculate the median of available n-pair covariance
    estimates at closes in the preceding 252-session training interval;
    freeze that threshold for the block.
  At decision close t, calculate x_t = adjusted_close_t /
    adjusted_close_(t−20) − 1.
  Calculate C_t(n), the centered sample covariance of the n most recent
    completed (x_u, close_(u+1)/close_u − 1) pairs, u = t−n,...,t−1.
  Target 100% SPY if x_t > 0 and C_t(n) exceeds the frozen threshold;
    otherwise target cash. Skip a gated fold with insufficient history.
  Compare with a date-matched x_t > 0 long-or-cash baseline without the gate.

The specification calls for changes at the observed adjusted open of t+1 and open-to-open holding returns. The supplied code says only one window produces capital-backed trades; it does not identify which window produced the supplied metrics.