Anyone asked to fund or replicate an LLM trading agent should begin with the counts from Xia and co-authors. They mapped 77 agentic-trading studies. Only 19 produced an action and completed the evaluation loop. Within that 19, two used time-consistent splits, one specified an explicit transaction-cost model, one documented its universe and survivorship treatment, and none supplied a fully replayable record. Those figures provide the most useful prior in this review.
Zhu and Cai's filter
Zhu and Cai review the literature without running a strategy. They provide no dataset, backtest or performance figure of their own. Their literature cutoff is 31 August 2026. The coverage includes listed common stocks and ETFs, centralized crypto spot, crypto perpetual futures, and on-chain decentralized exchange (DEX) venues. Fixed income, FX, options, commodities and private assets receive no full treatment.
The framework follows a five-stage chain. It begins with point-in-time information, then moves through representation and signal, portfolio and risk, orders and fills, and finally net alpha with capacity and persistence. Seven dimensions govern the assessment of each economic claim: temporality, selection control, portfolio mapping, implementation realism, risk and benchmark, external validity, operational provenance. The paper also classifies claims according to which of eight evidence objects dominates, spanning architecture capability through audited live capital. Zhu and Cai insist that these objects are not ranked levels. Unreported items are coded unknown; favourable treatment requires reported evidence.
Progress appears mainly upstream, in nonlinear cross-sectional prediction, text extraction, cost-aware portfolio learning and workflow integration. The harder profitability claims remain downstream. They require timestamped commitment, real fills, complete costs, capacity and persistence. This distance drives the paper's verdict, stated flatly in the abstract: within the examined public record, no general AI architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha.
When held-out data still cannot be traded
Temporality produces the review's strongest section because its examples alter the claimed economics. Zhang, Zhu and Linnainmaa trace a reported monthly ML alpha to look-ahead predictor alignment. Once the predictors are aligned with their actual availability, the alpha disappears. The Review of Financial Studies later issued an expression of concern on the original article (39(5):1555, 2026).
Lopez-Lira, Tang and Zhu show that LLMs can recall pre-cutoff financial values, and identifier masking does not provide a sufficient defence. Kelly, Malamud, Schwab and Xu use chronological model checkpoints instead. Zhu and Cai regard that method as the preferred fix, with genuinely post-cutoff prospective tests as the fallback.
The less obvious problem concerns numbers. Zhu and Cai argue that numerical pretraining contamination is harder to audit than textual contamination because time-series corpora may contain transformed or duplicated benchmark sequences with obscure provenance. Fine-tuning on a pre-cutoff slice cannot erase information already encoded by the base model. Anyone putting TimesFM or Chronos into a production forecast therefore needs a checkpoint whose entire training dataset ends before the trading period. I do not know of a widely used public one that documents this.
Results from foundation models give little comfort. On rolling-origin tests across five liquid US equities, Alonso and Franklin award pretrained time-series foundation models 8 of 10 task-level wins. Statistically significant gains over a random walk appear in only 2 model-asset comparisons. FinVerse evaluates 43 public foundation models on 116,897 financial series and finds that a strong generic forecasting rank need not produce a useful financial forecast. Tan and co-authors remove or replace the language component in LLM-based time-series methods and preserve or improve performance. Zhu and Cai carefully separate this result from TimesFM, Moirai or Chronos, which are trained on numerical sequences.
Costs choose the representation
The most usable modelling argument is simple: trading frictions influence which representation deserves to be learned. Jensen, Kelly, Malamud and Pedersen estimate the implementable efficient frontier directly. Their work shows how cost-agnostic prediction can favour fleeting opportunities with little scale.
Zhu and Cai avoid assigning a single direction to the magnitudes. Novy-Marx and Velikov show turnover wiping out anomaly profits. Frazzini, Israel and Moskowitz use live institutional trades and report lower costs and higher capacity than pessimistic academic models for some styles. The evidence licenses no blanket cost scalar in either direction.
Equation 1 lists the deductions: fees, spread, impact, delay, financing, venue. It assigns a magnitude to none of them. The framework supplies the fields and leaves calibration to the user.
Two disclosures govern our own work. We built a strategy from the paper's chain and ran it on listed US equities, ETFs and centralized crypto spot. The paper contains no backtest of its own, leaving nothing here to reproduce. Perpetual funding, on-chain cash flows, automated-market-maker inventory effects, gas auctions and transaction reordering are excluded rather than proxied because none survives substitution into spot or ETFs.
Our run also falls short of the paper's execution standard. We have no quote, bid-ask, order-book, fill, venue-fragmentation or gas data. Our costs are conservative assumptions, not realized microstructure.
Four markets, four P&L equations
Market stratification earns the paper its length. A perpetual position pays funding and faces liquidation as a discontinuity. A backtest using last-traded prices can therefore describe a path that the margined account would not have survived.
On-chain results bring a different P&L. Fritsch and Canidio find that arbitrage-related losses can exceed fee income in many large Uniswap pools. Capponi and Jia show arbitrage competition transferring much of the surplus into gas and infrastructure. In centralized spot, wash volume documented by Cong, Li, Tang and Yang and by Aloosh and Li contaminates any model trained on volume or order flow. The authors acknowledge that the AI-specific crypto evidence they examined is thinner than the equity evidence. Common criteria in these markets amount to a protocol rather than a body of results.
A verdict beyond external falsification
Zhu and Cai place "sustainable net alpha" above the target pursued by individual papers, and they acknowledge the choice. The definition "intentionally sets a higher bar". They also write, "Public insufficiency is not evidence of private absence." These concessions carry much of the argument because together they leave the headline conclusion untestable from outside.
The authors answer that a universal public algorithm would lose economic stability if inexpensive replication immediately crowded it. Their claim is therefore an epistemic bound: the evidence examined cannot support a conclusion stronger than its designs permit. That works as a judgement on research designs. The headline verdict remains beyond any outside test because the one record capable of refuting it is unpublished by construction.
Four load-bearing references are arXiv preprints: Alonso and Franklin, Lee and co-authors, Xia and co-authors, and Hua and co-authors. The representative-systems table also contains the first author's own DSA, an evidence-aware multi-agent orchestrator for multi-market stock research. Its entry reports conformance evidence only. There is no report-quality, forecast, return, or execution result.
Robinhood's roughly 100,000 opened Agentic Trading accounts and more than USD 100 million in custody measure adoption. Those figures are company-reported adoption, as the paper says.
The review does establish a gradient of evidence. Koijen and Levy report that optimized agentic systems increase explained contemporaneous return variation on a real-time earnings task from near 8% to close to 20%. Zhu and Cai refuse to classify this as a return, correctly so: contemporaneous explanation includes adjustment completed before a trader could be positioned.
Carlin, Israelsen and Wazzan assemble prospective daily LLM household portfolios. They find tilts toward large, growth, momentum and attention names, with no statistically significant abnormal return after their characteristic adjustment. Chen, Sialm and Xu report that AI hedge funds outperformed early and then lost some of that advantage. Return comovement among them was lower than the homogenization story predicts.
The next agent demo
According to the paper, the public record remains sparse on independently audited live-capital performance that persists across regimes and capital scales. A single study combining chronological checkpoints, a capacity curve over participation rate and a precommitted decision log would move me.
Until then, the Xia figures describe the reporting standard among the 77 agentic-trading papers Xia and co-authors coded, and nothing wider. They identify protocol and reporting gaps. They do not establish the profitability of every agentic system. Applied to the next agent demo placed in front of you, the filter costs nothing.