A clean equity curve can hide a strategy that trades the wrong rules. Wang, Wang and He find that even their strongest tested model fails silently on 27.5% of a 200-task subset.

A stop that moves when it should stay put

Ask for a stop anchored to the entry price. The generated script anchors it to the current close on every bar instead. It compiles, runs and plots a plausible curve. The authors found silent failures of this sort when they inspected the code. One breakout level uses a window containing the current bar, making a trigger impossible. Another error reads a higher-timeframe filter from the wrong candle. The open-weight models used htf.close[t // 4] to index the higher-timeframe series, reading a candle still in progress. Their engine catches the read; a conventional backtester would leak future data.

How MintEval makes the comparison

MintEval builds reference strategies from 6 signal, 2 filter, 3 sizing and 9 risk blocks, each on a discrete parameter grid. Each program combines one signal, a direction, a filter, a sizing rule and one to three risk layers. Some combinations require a breakout-then-retest state machine or a trailing stop that activates only after breakeven.

A third-vendor LLM turns each program into a 40 to 120 word desk message in trader slang (median 96 words, 11 slang hits). Gates require every parameter value to appear. A reader model from a fourth vendor tries to reconstruct the spec; 780 of 800 messages pass. The tested model receives the message and an interface document, then writes one strategy function.

Reference and generated functions run bar by bar on BTCUSDT 15-minute bars spanning about two years, both charged 6 bp per fill. ActionMatch measures how often their quantised targets agree on active bars, when either function holds or targets a position. Excluding flat bars keeps shared inactivity from inflating the score. A compiled run counts as a silent failure below 0.9 ActionMatch. The authors also report |ES|, the absolute difference in total return, alongside ΔSharpe and ΔMaxDD. They use an absolute gap because "an implementation error that happens to make money is still an error".

The 800 tasks are stratified jointly by τmax quintile and reference-return band. τ measures state span: when a stored value is read, it counts the bars since that key was last written. Its correlation with description length is -0.03, leaving long-memory tasks largely separate from long prompts.

The reference books have a median return of -35.1% after costs, and 8.9% make money. With zero friction, median return rises to -0.6% and 47.8% are profitable. Those losses reflect fees on high-turnover random rules. The authors argue that P&L sign does not drive the matching metric, and their results support that reading: references with worse returns are slightly easier to match (-0.11 ActionMatch per 100% of return, p=3e-05).

The average conceals the failures

The two closed low-cost models often compile code that misses the target. In the open setting, GPT-5.4-mini compiles 0.955 of tasks and scores 0.544 ActionMatch. It reproduces 0.087 exactly, with median |ES| of 442 bp. Its silent-failure share is 0.729; Claude Haiku's is 0.806.

On the 200-task subset, Claude Opus 5.5 compiles everything and scores 0.889 ActionMatch. GPT-5.4-mini scores 0.575 on those same tasks. Exact reproduction separates them more sharply: 0.575 for Opus and 0.085 for GPT-5.4-mini. Yet Opus fails silently on 0.275 of tasks. Under the paper's premise, one mistimed exit alters every subsequent decision. Its 0.889 average coexists with divergence on more than 10% of active bars for 27.5% of the subset.

The closed setting gives the models the block menu, separating their reading of the instructions from their coding. GPT-5.4-mini, Claude Haiku and Qwen-32B reach SpecMatch scores of 0.938, 0.946 and 0.935; Qwen-7B reaches 0.741. Among 1767 compiled runs with a fully correct spec, 0.792 still score below 0.9 ActionMatch (0.666 GPT-5.4-mini, 0.791 Haiku, 0.934 Qwen-32B). The menu moves GPT-5.4-mini from 0.544 to just 0.582. Persistent-state implementation accounts for much of the remaining gap, as the paper argues. Its abstract says models given the menu "identify the strategy almost perfectly".

Memory span accounts for some errors. In the open setting, GPT-5.4-mini drops from 0.674 to 0.473 across τmax quintiles. The pooled slope per decade of τ is -0.065 open and -0.088 closed. Adding block fixed effects and register count reduces those slopes to -0.041 (p=0.014) open and -0.058 (p=0.00023) closed. With full controls, the open estimate is -0.033 (p=0.049), which the authors call borderline. The closed estimate remains -0.063 (p=3.7e-05). Only the open estimate is borderline; the closed effect survives full controls. On the 200-task subset, Opus stays roughly flat, at 0.86 in the first quintile and 0.88 in the last. GPT-5.4-mini falls from 0.73 to 0.42 on the same tasks. The τ measure leaves Opus's remaining errors unexplained.

The judge passed the wrong implementations

The authors applied the QuantCode-Bench judge verbatim, using Claude Sonnet 4, to every compiled open-setting run it could judge from Opus and GPT-5.4-mini on the subset. It passed them all, including all 55 Opus silent failures and all 120 GPT-5.4-mini silent failures.

In a negative control, the judge rejected 7 of 7 pairings with code from a different signal family. The authors say this shows "the judge is not simply permissive". Seven controls provide thin support for that claim. The pass result stands on its own, though 26 GPT-5.4-mini runs went unjudged when the API budget ran out. As the authors put it, "It detects the wrong strategy; it cannot detect the right strategy implemented wrongly." QuantCode-Bench's pooled single-turn breakdown offers context: the judge rejects only 2.7% of attempts.

Would human instructions change the result?

These are synthetic tasks on one instrument, one window and one bar size. We did not find exact start and end dates. An LLM wrote the instructions, and the human baseline remains in progress. The authors disclose that Claude Opus 5.5, from the frontier model's own family, helped write the reference library and desk conventions.

The authors also call the 0.9 cutoff a convention. Changing it changes the headline: across models and settings, silent-failure shares range from 0.370 to 0.741 at 0.8 and from 0.372 to 0.884 at 0.99. For GPT-5.4-mini open, the shares are 0.655 at 0.8 and 0.840 at 0.99. Qwen-7B's low open-setting share of 0.372 comes with a compilation rate of only 0.374.

We cannot reproduce the leaderboard without the named model versions and original tasks. We also cannot trade BTCUSDT. Our version uses our platform's BTC price series, newly generated tasks and the models available to us. Reference and generated code run on identical bars with identical frictions, making this a test of the validation method. The paper's scores will not transfer to our run.

A planned real-strategy subset would change my view if it showed a frontier silent rate well under 0.275 on human-written intents.