RADAR's 1.062 Sharpe is too thin a basis for a trade without stated trading costs. It comes from about five test years. Roughly 0.2 of Sharpe belongs to the paper's new idea, added to an embedding model that already beat equal weight. That 0.2 warrants a test, not a position on this evidence alone.

The trade Koa, Li and Huang built

The study uses daily adjusted closes from Yahoo Finance and news from the FNSPID corpus. Its universe contains 377 US equities appearing in both the S&P 500 and FNSPID, from June 2011 to June 2020, about 2,270 trading days. A GRU with attention encodes 60 days of prices. A second GRU handles FinBERT embeddings of the news. A bilinear layer combines the two into a 64-dimensional market state.

The new idea concerns the noise fed to diffusion. Standard diffusion adds isotropic Gaussian noise to an embedding, then learns to remove it. RADAR builds a context bank instead. Each entry contains news from one date and the prices that follow until the next news item. For the current price and news embeddings, the model retrieves the ten closest entries by cosine similarity. It uses their mean and variance to set the noise distribution for 100 forward diffusion steps, after which a self-attention denoiser reconstructs the clean embeddings. Noise drawn from similar past regimes is meant to teach plausible variation for the current market state; isotropic noise has no dependence on that state.

The cleaned representation determines portfolio weights. Training uses the stochastic discount factor (SDF) objective of Kelly and coauthors: the second moment of 1 minus portfolio return, with an L2 weight penalty and a diffusion reconstruction loss. The book rebalances every 7 days. Evaluation advances a year at a time, with 4 years of training, a 90:10 validation split and then 1 year out-of-sample. The paper reports Sharpe 1.062, Sortino 1.352 and Calmar 0.924. Equal weight, its strongest baseline, has Sharpe 0.795; the S&P 500 has 0.513.

Gaussian noise takes Sharpe from 0.859 to 0.411

The ablation table puts the diffusion claim under pressure. On raw features, the Kelly-style SDF model scores 0.505 Sharpe. Learned embeddings lift that to 0.859, above equal weight already. Ordinary Gaussian diffusion then takes Sharpe down to 0.411. Maximum drawdown reaches 0.784, and volatility reaches 0.525.

The authors acknowledge the damage: "diffusion noise appears to act more as distortion than regularization." They deserve credit for showing it. Their case for retrieval is that the context bank makes noise "grounded in meaningful past market conditions". Retrieval has to rescue a diffusion step that otherwise hurts. Against the embeddings-only model, full RADAR adds 0.203 Sharpe (1.062 against 0.859). The 0.203 measures the paper's contribution. The abstract claims the best risk-adjusted performance without mentioning costs. RATD, another retrieval-plus-diffusion model in the comparison, records 0.680, last among the forecasting group.

The news-only ablation is revealing too. Remove prices and Sharpe is 0.797, with cumulative return 1.003 and volatility 0.200. Equal weight records 0.795, 1.005 and 0.200. The news-only book has effectively converged to 1/N; the active bets come from prices. Yet prices alone yield 0.650 Sharpe, 0.332 volatility and a 0.573 drawdown. The edge needs both channels.

Is 0.285 a year worth 0.271 volatility?

The paper openly gives the ratios priority over the other metrics. RADAR wins on ratios. Its annualized return is 0.285 against 0.149 for equal weight, while volatility is 0.271 against 0.200. Maximum drawdown favours RADAR, 0.308 against 0.380. That drawdown gap is the paper's strongest evidence in its favour.

The risk profile raises questions for a large-cap book. CAPM beta is 0.977, yet R² is only 0.476: the market leaves more than half the daily variance of a portfolio drawn from 377 S&P names unexplained. Near-one beta alongside low R² points toward concentrated or long-short positions. We found an L2 penalty on weights, with no budget or long-only constraint, leaving gross exposure and leverage unknown. We also found no turnover or transaction-cost figures, despite rebalancing every 7 days.

CAPM alpha is 18.71% annualized (t = 2.159, p = 0.016). Under Fama-French three-factor it falls to 15.47% (t = 1.814, p = 0.035), with a value (HML) loading of -0.173. The regression p-values are one-sided, as the t-stats imply. The paper specifies a one-sided hypothesis only for its return-difference tests. A t of 1.81 gives the 15% alpha a weak footing.

Five test years, and a retrieval window to audit

The significance tables appear forceful. Against equal weight, the Ledoit-Wolf Sharpe-difference p is 0.0079; against iTransformer, the weakest pooled Ledoit-Wolf p is 0.0402. The bootstrap Sharpe test against SDF gives 0.0437. Those tests pool five seeds of one strategy over one market path. Fisher aggregation across seeds addresses seed luck. It supplies no independent history. The history comprises about five annual test windows. Their dates are unlisted, although the 4-plus-1-year roll within June 2011 to June 2020 implies the last five years.

Three procedural questions matter more than those p-values.

  1. Retrieval permits any segment with a news date before t. A segment's prices continue until the day before the next news event, so the latest permitted segment can, as written, include prices after t. We did not find a truncation rule.
  2. We did not find whether S&P 500 membership was point-in-time. A single set of 377 names appearing in both the S&P 500 and FNSPID across 2011 to 2020 would introduce survivorship if selected using the full window.
  3. The top-K sweep selecting K = 10 reports Sharpe across seeds. We could not tell whether it used validation or test windows.

The comparison group is weak as well. BSV, DKKM, LinAttn and the SDF transformer return 0.036 to 0.046 annualized, around half the S&P's 0.083. They take much less risk: SDF transformer volatility is 0.083, against 0.191 for the index. The authors say the ordering agrees with the original SDF work. Their Sharpes of 0.361 to 0.505 all fall short of the index's 0.513. Baselines trailing the index leave an easy hurdle.

I would change my mind if retrieval were truncated at t and RADAR still clearly beat the 0.859 embeddings-only variant with a point-in-time universe and realistic costs. That would make the retrieved-noise idea worth keeping.

If the gap to 0.859 vanishes, the paper has found a good encoder and a complicated way to avoid damaging it.