AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

This quantum StatArb edge rests on 2020 and zero costs

GBS Roots beats SPONGE in one in-sample year, but a classical sampler tracks it and no device runs it.

2026-07-27 · 4 min read

Reviewing: Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Top-500 Large-Cap Residual-Correlation Cluster StatArb with GBS-Roots-Inspired Dense-Subgraph Clustering

Backtest period 2015-01-01 to 2024-12-31 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper

  • The paper's method needs quantum hardware or a quantum sampler, which we do not have. Any strategy we build is a substitute (typically the paper's own classical baseline), so our backtest does not test the paper's claim.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (18)

  • R^res_{t,i} := R_{t,i} − β_i R_{mkt,t}, where β_i := Cov(R_i,R_mkt) / Var(R_mkt).: (invalidates: Paper Table I; Paper Table II; Paper Welch test; Paper regime result; Paper volatility relationship)
  • C_ij := [∑_{t=T−w}^{T−1} (R^res_{t,i} − R̄^res_i)(R^res_{t,j} − R̄^res_j)] / [(w−1)σ_i σ_j].: (invalidates: Paper Table I; Paper Table II; Paper Welch test; Paper structural result; Paper regime result)
  • C = 1/(w−1) X^T X.: (invalidates: Paper Table I; Paper Table II; Paper Welch test)
  • j_i = PW, if ∑_{t=T−w}^{T−1} Δ_{t,j_i} > p; j_i = PL, if ∑_{t=T−w}^{T−1} Δ_{t,j_i} < −p.: (invalidates: Paper Table I; Paper Table II; Paper Welch test)

14 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2015Total -5.8%Dec 2024
Sharpe
-1.00
Total Return
-5.8%
Max Drawdown
-6.2%
CAGR
-0.6%
Volatility
0.6%
Trades
38,366

What the paper reports for its own strategy

  • GBS Roots (lossless 100-stock S&P 500, year 2020, zero transaction costs assumed): total return 0.236±0.031, Sharpe 2.248±0.293, Sortino 3.708±0.673 (annualized, ~20 runs)
  • GBS Boost (lossless 100-stock S&P 500, year 2020, zero transaction costs assumed): total return 0.239±0.035, Sharpe 2.050±0.190, Sortino 3.188±0.521 (~20 runs)
  • GBS methods beat SPONGE in 85% of cases at photon loss ≤40% (50-stock universe, 2020, with displacement)

Start with what a desk would actually trade: a dollar-neutral contrarian book that shorts recent winners and longs recent losers inside clusters of co-moving S&P 500 stocks, with the clusters found by a Gaussian Boson Sampling heuristic instead of spectral methods. The paper's claim is narrow and specific. On a lossless 100-stock universe over 2020, the novel GBS Roots clustering produced total return 0.236 (±0.031) and Sharpe 2.248 (±0.293), edging SPONGE's 0.214 return (Sharpe 2.242, ±0.669) and Spectral's 0.200 (Sharpe 2.035). A two-tailed Welch's t-test puts the return gap over SPONGE at p=0.023, while the Sharpe gap is a coin flip (p=0.949).

Before any number of ours appears, the honest disclosure: we do not have a photonic device or a quantum sampler, so we could not run the method the paper is actually about. What we built is the classical, deterministic GBS-Roots-inspired dense-subgraph heuristic, the same family of clustering the authors themselves benchmark against. Whatever that build does, it measures a classical clustering rule on a different universe and decade, not the quantum sampling the paper credits with the edge. We ran it on annual top-500 US large caps from 2015 to 2024, residualizing against SPY plus sector ETFs, charging four tenths of a cent a share with no slippage in the primary pass. Our one automated pass did not surface a headline Sharpe we would stand behind, and because it is a machine-built substitute rather than a replication the authors would sign, read that as a statement about our implementation and universe, not a verdict on their result. We simply cannot test their claim without the sampler.

What the paper is actually measuring

The economic headline lives almost entirely in one volatile year. Main simulations run over 2020, with regime checks on 2008, 2017 and 2022, and the returns are gross: the discussion states plainly that trades "do not incur price impact and there are zero transaction costs." That matters because the strategy rolls the window forward every three days, holds three days, and reshuffles winners and losers within clusters. Quantum methods showed much lower Jaccard similarity across windows (0.163 for GBS Roots versus 0.344 for Spectral), which the authors read as more dynamic clusters. More dynamic clusters mean more turnover, and turnover is exactly what a zero-cost assumption hides.

GBS Roots' real selling point is the spread across runs. It matched SPONGE's Sharpe with less than half the cross-run standard deviation (±0.293 versus ±0.669) across roughly 20 runs, against about 100 for the classical methods. Steadier output from a stochastic sampler is a real property. Whether it survives 100 more runs on a different year is untested here.

The hardware question swallows the result

Everything rests on low photon loss. In the 50-stock tests, GBS methods beat SPONGE in 85% of cases at loss at or below 40%, but without coherent displacement they fall below both Spectral and SPONGE once loss passes 70%, decaying toward the random baseline. Displacement rescues this, and the authors are candid that the main results come from simulated GBS on classical hardware, not a physical device. So the finding is really a statement about a hardware regime that does not yet exist at the scale needed (the classical hafnian cost is O(N^3 2^N), which is why the universe is capped at 100).

The part that should bother a quant most: QIC-GBS, a purely classical proxy, reproduces the GBS Boost dynamics efficiently. If a classical engine already samples the same distribution, the case for the photonic device rests on GBS Roots specifically, and on loss staying low. The authors flag another asymmetry themselves, and it cuts against them: the classical benchmarks are handed a market-informed cluster count K, while the quantum methods choose their own. Handing the classical side the exact number of clusters looks more like help than a penalty.

Density is not the signal

The cleanest result in the paper is a negative one. GBS Boost produced by far the highest cluster value (V=17.031) yet only 0.239 return, while GBS Roots reached 0.236 with V=3.945. SPONGE had the highest weighted density (0.336) and still trailed on return. Under an idealized model the return-maximizing weighted density is 2/3, not 1. Maximizing the graph objective the sampler is good at does not maximize PnL, which means the quantum advantage at dense-subgraph search is aimed at the wrong target. If tighter clusters offer fewer mean-reversion opportunities, then the thing GBS does natively is not the thing the strategy needs.

Grant the authors a genuine, statistically supported in-sample return edge for GBS Roots over SPONGE in one high-volatility year, and a real variance reduction. What would change my mind is a multi-year test with realistic costs and turnover, run on physical hardware or at least showing the classical QIC-GBS proxy cannot match GBS Roots. Until then this reads as a careful correlation-clustering study wearing a quantum badge that the economics do not yet earn.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each rebalance date T:
  1. Select annual top-500 US stocks by capitalization.
  2. Remove secondary share classes, preferred/tracking/baby-bond-like instruments, and near-duplicates with trailing residual correlation &gt; 0.98.
  3. Build dividend-adjusted close-to-close returns.
  4. Estimate 252-day OLS residuals versus SPY plus the stock sector ETF when available; fall back to SPY-only when sector data are missing.
  5. Form a 252-day residual-return correlation matrix C.
  6. Apply RMT cleaning: nullify market mode, drop Marchenko-Pastur noise bulk, retain significant PCs above lambda+.
  7. Cluster the cleaned positive residual-correlation graph using the deterministic GBS-Roots-inspired greedy dense-subgraph heuristic.
  8. For each cluster and formation window w in {5, 10, 20}:
       delta_i,t = residual_return_i,t - mean residual_return of valid cluster members at t
       score_i = sum(delta_i,t over the last w days)
       PL = names with score_i &lt; 0; PW = names with score_i &gt; 0
       skip clusters without at least one PL and one PW
  9. Portfolio per eligible cluster:
       long PL uniformly with 0.5 gross per cluster
       short PW uniformly with 0.5 gross per cluster
       keep dollar-neutral; optionally adjust beta to SPY if feasible
 10. Execute at close in the backtest, hold for 5, 10, or 20 trading days as separate non-overlapping sleeves.
 11. If an open sleeve reaches +2% cumulative return before scheduled exit, liquidate at the next available close.