Randomly reassign this LLM agent's 310 action changes across eligible earnings events, and they earn +25.8 bps/event. The agent earned +15.2. Its event picks fell 24.6 bps per intervention short of the matched benchmark, a result at the 32nd percentile (p=0.63). The shortfall is statistically undetectable. The authors put it plainly in the abstract: "The agent does not detectably outperform matched random assignments."
Chen, Xu and colleagues built a three-role agent on Llama-3.3-70B. A Scout turns X and Reddit posts into firm notes; a Strategist chooses long, short or flat; a Reviewer checks the output. Beneath it sits an elastic-net forecaster with an out-of-sample IC of +0.006. Social data are meant to supply the edge. Remove that channel, and the system follows the rule on 723 of 723 events, so every action change comes from it. The frozen threshold rule goes long above one forecast standard deviation (0.00995), short below the negative threshold, and flat otherwise. The agent starts with that position and can change it. It did so on 310 of 723 events.
The sample covers 723 quarterly announcements from 44 U.S. consumer-facing firms over 2020 to 2024. Returns are measured as a three-day Carhart four-factor CAR[-1,+1]. The authors' Agent Policy-Value Audit separates the payoff from the agent's trade mix from the payoff from choosing events. They test it on semi-synthetic no-skill agents and on their own agent's 723 events. Deployment value averages (agent position minus rule position) times the CAR; agreement with the rule scores zero. For the composition benchmark, they preserve the count of each ordered action change, such as flat-to-long, and randomly place those changes among events sharing the rule's action. That gives the expected payoff from the observed trade mix without event picking. Selection value subtracts this benchmark from deployment value and is tested against 5,000 reassignment draws.
Why random flat-to-long trades earn +67 bps
The rule was flat on 547 events where the agent faced a genuine decision. It changed 294 of those positions to long. Reactions in that flat pool averaged +67 bps, with 52.3% positive, so a randomly assigned flat-to-long change earned +67.4 bps on average. The pattern predates the test window: among 285 pre-2020 events, the flat pool had a +69 bps mean reaction and the same 0.523 base accuracy.
A paired test against zero awards the +67 to the agent. With real returns and the real transition mix but no skill, the zero-centered paired test flags skill in 11.6% of 5,000 exposure-confound replications. The matched audit flags 5.3%. A separate 1,000-replication size test puts the firm-clustered paired test at 17.1%. Across a 64-cell stress grid, the zero-centered rate reaches.88; the audit remains between.04 and.07. "Larger samples make this problem worse, not better," the authors write. The price is power:.783 at directional accuracy 0.60, versus.846 to.917 for the procedures that over-reject.
They find another failure under decision-time overlap. Transition matching by itself raises false attribution from.592 to.648. Restricting the score to returns after the decision brings it to.000, which is why the audit requires future-only scoring.
The agent's event picks
The agent made +35.5 bps per intervention; the matched-null mean was +60.1. Selection value was therefore -24.6 bps per intervention, at the 32nd percentile of reassignments (p=0.63). Its 294 flat-to-long changes earned +47.7 against +67.4. The eight "other" changes, four long-to-flat, two short-to-flat and two long-to-short, averaged -522.8 bps per intervention. Wins came on 155/310 changes, or 0.500, against a no-skill base of 0.521; that win rate sits at the 13th percentile. Per event, +15.2 bps of gross deployment value consists of +25.8 from composition and -10.6 from selection.
The authors treat this as attribution and stop short of claiming equivalence. The gross interval spans [-45.2, +69.2] bps/event. The design reaches 80% power only around 142 bps per intervention, leaving small selection effects hard to see while weighing against large skill. The gross gain also depends heavily on thirteen interventions in five cells excluded by the firm-preserving null. Without those interventions, the remainder averages 8.1 bps, versus 35.5 overall. Selection value remains negative across all 44 leave-one-firm-out reruns.
The social feed changes how often the agent trades; the placebos find no measurable change in trade payoff.
Give it a wrong-firm feed and the intervention rate drops from 0.463 to 0.029. Every return-contrast interval includes zero.
Costs erase the gross gain
The authors argue for reporting deployment and selection value separately. They allow that positive deployment value with little selection value "may still justify deployment net of cost." At 40 bps per nonzero position, this agent's +15.2 bps/event gross becomes -1.2 net. The interval is [-59.8, +54.3], with a one-sided p of 0.52. Break-even comes to 37.1 bps per net additional position across 296 such positions. Under the switching-cost convention, the 20, 40 and 60 bps overlays produce +6.6, -2.1 and -10.7 bps/event. Scoring uses factor-adjusted abnormal returns, and the authors assume negligible market impact.
The 166 events after Llama's December 2023 cutoff look stronger: +78.9 bps/event gross (interval [-38.6, +202.5]) and +62.5 net. Per-intervention payoff is 192.6 against a matched 134.5, at the 68th percentile. The authors regard that as consistent with sampling noise given the interval's width, and I agree. The result rests on 68 changes.
What the audit leaves open
The authors identify two limits. By construction, skill in intervention frequency or transition choice enters the composition benchmark. A policy that knows how often to go long, or knows flat-to-long beats flat-to-short in this pool, receives no selection credit for that skill. The audit also assumes within-pool exchangeability. When no-skill assignment depends on return-shifting covariates, whether volatility, size, social volume or a nonlinear mix, rejection falls to.013 to.020. The authors call this a sensitivity diagnostic. Their analysis is retrospective: they devised the decomposition after freezing the traces and left the 2025 events unused.
We cannot rerun the authors' agent. We do not hold its frozen decisions or the bespoke universe filters for analyst coverage and social footprint. We are auditing a policy we constructed locally on earnings events and have no result to report yet. Any result from our run will address whether the audit is practical to build; it will say nothing about this agent's reported figures. Rebuilding it calls for a frozen baseline with logged actions, pools keyed to that baseline's action (optionally crossed with year or firm), and a scoring window beginning after the decision. The authors freeze in-house inputs at the close of t-2 for exactly this reason. The benchmark mean itself requires no simulation: it is the count-weighted pool average.
A held-out run in which per-intervention payoff exceeds the matched mean by more than the roughly 142 bps this design can detect would change my view of the agent.