AlphaQCM's 1.85 net long-short Sharpe in the S&P 100 is worth investigating. The fixed constituent list leaves its tradability in doubt.
The search Wang and Ventre ran
Factor mining searches for formulas, or Python code, that score stocks using daily price and volume. Those scores target the next 20 trading days' return. Portfolios buy the highest-ranked stocks and short the lowest-ranked ones.
Wang and Ventre ran eight published miners alongside a genetic-programming baseline, leaving each method to use its native search. The nine comprise that baseline, three reinforcement-learning or guided-search methods (AlphaGen, AlphaQCM, AlphaCFG), two generative searchers (AlphaForge, AlphaSAGE) and three LLM agents on Qwen3.8-27B. Across five markets and three seeds, the grid produced 135 runs. It found 4,914 factors; 4,698 (95.6%) yielded usable signals.
Yahoo Finance daily OHLCV supplies fixed constituent lists for the S&P 100, Hang Seng, CSI 300, FTSE 100 and Nikkei 225, totaling 782 stocks. Training covers 2017 to 2021, validation 2022 to 2023, and testing January 2024 to October 2025. The authors rank each pool by validation IC, discard factors whose correlation with a kept factor exceeds 0.8, then combine the survivors with ridge regression. Each composite trades a long-only top quintile and a dollar-neutral book with the top quintile long and bottom quintile short. The hand-built reference consists of 71 implementable Alpha101 formulas.
The main finding is an absence of consistent dominance by any mining paradigm. I buy that finding. Long-only net Sharpe ranges from 1.027 to 1.313; Alpha101 comes third at 1.241. AlphaQCM's long-short Sharpe varies across markets with a standard deviation of 1.032, above its pooled mean of 0.668. Market variation this large makes the ranking null credible. The authors confine their conclusion to "within this benchmark, rather than a universal ranking." For a trader, the question becomes whether a particular result merits a rebuild.
Why keep AlphaQCM on the screen?
In the S&P 100 long-short results, AlphaQCM reaches 1.85±0.30 across three seeds. GP and R&D-Agent each reach 1.59, with spreads of ±0.24 and ±0.08 respectively. AlphaGen records 1.46±0.34, Alpha101 0.97, and AlphaAgent 0.79±0.62.
Every method makes money in this S&P 100 test. Even the weakest miner and a 2016 formula library earn positive Sharpe on those same 100 names in 2024 and 2025. AlphaQCM leads GP by 0.26 of Sharpe, roughly one seed standard deviation. Its other claim is trading frequency: pooled across five markets, AlphaQCM turns over its long-short book 4.58x annually, versus 9.81x for AlphaAgent. It also remains positive after the paper's pooled neutralization, unlike the other automated miners.
Costs against a thin return
The paper rebalances every 20 trading days and charges 10bp per dollar traded, plus short borrow. We did not find the borrow rate in the main text. Its cost sensitivity reaches 50bp, where the authors describe long-short profitability as much more fragile than long-only: most methods approach or drop below zero as costs increase. Net CAGR across methods starts at 0.3% and tops out at 6.5%. The authors also acknowledge that market impact and capacity are unmodeled.
The test lasts 22 months. With 20-day holds, that amounts to roughly two dozen non-overlapping periods by our count, and we did not find a significance test for the portfolio Sharpes. GP illustrates how little the ranking statistic alone settles: its test composite has a RankIC of -0.0058 and still earns a long-short Sharpe of 0.340. The quintile tails can pay even when the full ordering fails. AlphaQCM's RankIC is 0.0056.
Today's constituents, yesterday's trades
The authors hold constituent sets fixed rather than reconstructing historical membership. They wanted a consistent cross-section and acknowledge that the choice "may introduce survivorship bias." Point-in-time investment universes appear in their conclusion as future work. Their S&P 100 list contains PLTR, GEV and UBER, names that reached a top-100 list because of where they ended up. Training begins in 2017, while every US backtest uses a list drawn from current index membership. That selection favors survivors.
A shared universe helps compare methods. It cannot validate an absolute US Sharpe.
The same bias plausibly affects all nine methods, leaving the paper's ranking null on firmer ground than its absolute returns. The 1.85 Sharpe remains exposed.
The missing US neutralized number
Pooled across five markets, AlphaQCM has the strongest neutralized case among the automated miners. After controls for beta, volatility, momentum, reversal and liquidity, its composite retains a test residual IC of 0.0175, against 0.0191 raw. Neutralized long-short CAGR drops from 3.7% to 1.3%; Sharpe falls from 0.668 to 0.283. Turnover climbs from 4.58x to 6.02x. R&D-Agent fares worse: CAGR goes from 6.5% to -1.6%, and Sharpe from 0.679 to -0.449.
AlphaQCM is the only automated method left positive. Alpha101 does better on the long-short book, with Sharpe rising from 0.319 to 0.540 and exceeding every miner. These controls cover five price-and-volume styles, with no industries. The tables provide no US-only neutralized Sharpe. Figure 17 shows the US neutralized long-short wealth path without printing a figure for Sharpe, so a rebuild must calculate it.
What our run can establish
We cannot trade the paper's book. Our data cover US large caps; we cannot test its five-market comparison or reproduce exact point-in-time index membership. We are also not running the nine miners. Instead, our run uses a bounded symbolic-expression search and fixed formula baselines under the paper's evaluation protocol: the same splits, 0.8 filter, ridge, quintile books, 20-day rebalance and 10bp. This is a test of the evaluation protocol, with results pending. It says nothing about AlphaQCM's standing against the LLM agents.
AlphaQCM's US cell would deserve a closer look if its long-short Sharpe beat Alpha101 after neutralization on a point-in-time universe. The comparison is demanding. Across five markets, Alpha101's long-short Sharpe rose from 0.319 to 0.540 under the paper's neutralization; its US long-short Sharpe is already 0.97 before neutralization.