AlphaPADI's S&P 500 RankIC warrants a replication. Its portfolio Sharpes demand an explanation first. The paper tests pool context before adding rewards, then leaves it in the full model. The final result therefore cannot separate the contribution of pool-conditioned generation from reward-ranked selection and preference training.

Generating the whole pool

Formulaic alphas trade as a pool combined by one model. A formula earns its place through what it adds to that pool. AlphaGen (PPO) and AlphaSAGE (a GFlowNet) score formulas with the pool in mind, yet each generates one formula at a time. AlphaPADI generates the whole pool instead.

Each candidate contains K formulas, with K set to 20 or 50. The formulas use Reverse Polish Notation and a vocabulary of 61 tokens: six price-volume fields, 14 constants, five windows from 1 to 40 days and 36 operators. A formula can contain at most 20 tokens. A discrete diffusion model masks a leaf, a subtree or an entire formula in the syntax tree. A small Transformer with hidden size 128 refills the masks one at a time while seeing the other pool members. If two members already contain moving-average signals, the generator can rewrite that subtree in the third.

The pool reward starts with the training-period mean daily Spearman of a ridge combiner fit on ranked members. A second term penalizes members whose maximum absolute correlation with another member exceeds a tolerance. High-scoring candidates go into an elite buffer of 128 pools. Training draws on that buffer, combining reconstruction loss with a DPO-style preference loss on reward-ordered pool pairs.

The authors use daily Qlib bars for CSI300, CSI500 and S&P 500. The target is the 20-day forward return; training covers 2016-2020, validation 2021-2022 and test 2023-2025. Non-LLM miners each receive 40,000 formula evaluations. Every pool uses the same ridge combiner, fitted on training data and then frozen. Across five runs, AlphaPADI ranks first on all seven metrics in all three markets.

How much of the S&P 500 result trades?

The test RankIC is 0.0452 for AlphaPADI and 0.0220 for both AlphaForge and AlphaSAGE, a gain the authors report as 105.5%. IC reaches 0.0515 against AlphaSAGE's 0.0319. RankIC margins in the Chinese markets are smaller, at 21.1% and 27.0%. The paper takes the absolute value of mean daily correlation for both IC and RankIC, leaving the sign absent from those columns.

The portfolio backtest buys an equal-weighted top 20% each day and assigns the cohort 1/20 of fixed capital. It holds each cohort for 20 trading days, charging 5bp on buys and 15bp on sells. With twenty cohorts live at 1/20 apiece, this is a fully invested long-only book.

On S&P 500, AlphaPADI reports a 20.61% simple annual return and a Sharpe of 5.84. For that fully invested long equity book, the figures imply roughly 3.5% annualized volatility. The static Alpha158 library reaches an S&P 500 Sharpe of 3.86 under the same protocol. Every method's S&P 500 Sharpe falls between 3.86 and 5.84. A fixed technical library approaching a Sharpe of 4 long-only makes the accounting behind that level worth examining. I could not find a volatility-compressing step in the described procedure, so this remains an observation.

The trading edge is much narrower than the RankIC gap: Sharpe is 5.84 against AlphaForge's 5.39, with drawdown of 8.60% against 9.52%.

The missing CSI300 row

The component study is cumulative and confined to CSI300. Diffusion without pool context starts at IC 0.0215 and RankIC 0.0334. Adding pool context raises RankIC to 0.0386 and ICIR from 0.1434 to 0.2263. Yet annual return slips from 7.80% to 7.31%, and Sharpe from 1.97 to 1.90. Drawdown improves from 27.08% to 22.49%. The authors write that "pool-level rewards are needed to turn pool-aware reconstruction into trading gains."

Those rewards arrive later. The predictive reward moves RankIC to 0.0463. The diversity reward brings it to 0.0531 and reduces drawdown from 23.33% to 19.15%. Preference loss produces the largest single RankIC step, to 0.0659.

Pool context is first measured against a model that buffers its latest candidates and samples them uniformly, before either reward is active. The table never removes pool context while retaining both rewards and preference loss. A pool-conditioned generator is unnecessary for pool rewards and pairwise preference to operate. The missing row leaves the generator's share of the final 0.0659 RankIC unresolved. The authors do not repeat this study on S&P 500, where the reported margin is largest.

At the end of search, training-split IC on CSI300 is 0.1245 against 0.0439 in test. AlphaPADI's learning rate, diversity weight and reconstruction weight are tuned on validation; the baselines use open-source defaults, with only K selected.

The frozen combiner also matters for comparison. The authors say it "cannot compensate for factor decay over the three-year test period, which makes the evaluation stricter." That reasoning applies more cleanly to AlphaPADI than to AlphaForge, which was designed for dynamic combination. In the paper's results, second-best entries are spread across six baselines.

Replication on S&P 500

We are building a U.S. replication using their settings: daily bars, point-in-time constituents, and the 2016-2020, 2021-2022 and 2023-2025 split. The budget is 40,000 distinct formula evaluations. Validation selects K from 20 and 50. We choose the ridge penalty on 2016-2019 against 2020 with a 20-day gap, then freeze the combiner.

The test targets are RankIC above 0.0220 and IC above 0.0319. Baselines must use the same top-20%, 20-day cohort rules and costs.

A full-model run without pool context would change my mind. It would retain both rewards and preference loss. If its RankIC lands near 0.0452, the selection loop accounts for the result. If the pool-aware version beats it by a margin comparable to the 0.0052 added by pool context on CSI300 (0.0334 to 0.0386), the claim for pool-conditioned generation gains support.

For now, AlphaPADI looks like a reward-guided pool search that doubles the best baseline's S&P 500 RankIC, with a diffusion generator attached. The abstract calls the results "validating pool-aware generation as an effective framework." The conclusion limits its claim to findings that "support pool-level search under the evaluated settings," which is what the ablation establishes.