Start with what a desk would actually trade: a dollar-neutral contrarian book that shorts recent winners and longs recent losers inside clusters of co-moving S&P 500 stocks, with the clusters found by a Gaussian Boson Sampling heuristic instead of spectral methods. The paper's claim is narrow and specific. On a lossless 100-stock universe over 2020, the novel GBS Roots clustering produced total return 0.236 (±0.031) and Sharpe 2.248 (±0.293), edging SPONGE's 0.214 return (Sharpe 2.242, ±0.669) and Spectral's 0.200 (Sharpe 2.035). A two-tailed Welch's t-test puts the return gap over SPONGE at p=0.023, while the Sharpe gap is a coin flip (p=0.949).
Before any number of ours appears, the honest disclosure: we do not have a photonic device or a quantum sampler, so we could not run the method the paper is actually about. What we built is the classical, deterministic GBS-Roots-inspired dense-subgraph heuristic, the same family of clustering the authors themselves benchmark against. Whatever that build does, it measures a classical clustering rule on a different universe and decade, not the quantum sampling the paper credits with the edge. We ran it on annual top-500 US large caps from 2015 to 2024, residualizing against SPY plus sector ETFs, charging four tenths of a cent a share with no slippage in the primary pass. Our one automated pass did not surface a headline Sharpe we would stand behind, and because it is a machine-built substitute rather than a replication the authors would sign, read that as a statement about our implementation and universe, not a verdict on their result. We simply cannot test their claim without the sampler.
What the paper is actually measuring
The economic headline lives almost entirely in one volatile year. Main simulations run over 2020, with regime checks on 2008, 2017 and 2022, and the returns are gross: the discussion states plainly that trades "do not incur price impact and there are zero transaction costs." That matters because the strategy rolls the window forward every three days, holds three days, and reshuffles winners and losers within clusters. Quantum methods showed much lower Jaccard similarity across windows (0.163 for GBS Roots versus 0.344 for Spectral), which the authors read as more dynamic clusters. More dynamic clusters mean more turnover, and turnover is exactly what a zero-cost assumption hides.
GBS Roots' real selling point is the spread across runs. It matched SPONGE's Sharpe with less than half the cross-run standard deviation (±0.293 versus ±0.669) across roughly 20 runs, against about 100 for the classical methods. Steadier output from a stochastic sampler is a real property. Whether it survives 100 more runs on a different year is untested here.
The hardware question swallows the result
Everything rests on low photon loss. In the 50-stock tests, GBS methods beat SPONGE in 85% of cases at loss at or below 40%, but without coherent displacement they fall below both Spectral and SPONGE once loss passes 70%, decaying toward the random baseline. Displacement rescues this, and the authors are candid that the main results come from simulated GBS on classical hardware, not a physical device. So the finding is really a statement about a hardware regime that does not yet exist at the scale needed (the classical hafnian cost is O(N^3 2^N), which is why the universe is capped at 100).
The part that should bother a quant most: QIC-GBS, a purely classical proxy, reproduces the GBS Boost dynamics efficiently. If a classical engine already samples the same distribution, the case for the photonic device rests on GBS Roots specifically, and on loss staying low. The authors flag another asymmetry themselves, and it cuts against them: the classical benchmarks are handed a market-informed cluster count K, while the quantum methods choose their own. Handing the classical side the exact number of clusters looks more like help than a penalty.
Density is not the signal
The cleanest result in the paper is a negative one. GBS Boost produced by far the highest cluster value (V=17.031) yet only 0.239 return, while GBS Roots reached 0.236 with V=3.945. SPONGE had the highest weighted density (0.336) and still trailed on return. Under an idealized model the return-maximizing weighted density is 2/3, not 1. Maximizing the graph objective the sampler is good at does not maximize PnL, which means the quantum advantage at dense-subgraph search is aimed at the wrong target. If tighter clusters offer fewer mean-reversion opportunities, then the thing GBS does natively is not the thing the strategy needs.
Grant the authors a genuine, statistically supported in-sample return edge for GBS Roots over SPONGE in one high-volatility year, and a real variance reduction. What would change my mind is a multi-year test with realistic costs and turnover, run on physical hardware or at least showing the classical QIC-GBS proxy cannot match GBS Roots. Until then this reads as a careful correlation-clustering study wearing a quantum badge that the economics do not yet earn.