For a desk building research tools, about 800 validated examples improved a 35B model's Backtrader code more than 5 to 6 billion tokens of trading repositories. Chernysh, Ekhtibarov and Zmitrovich report the larger gain after fine-tuning, which followed continued pretraining. They also found a cost: specialization broke tool-call formatting, and recovery fine-tuning did not restore the base checkpoint's repository-level agent performance. The order of training limits what can be credited to either stage.
From a request to runnable code
A trader asks for a long entry when RSI crosses 30, with a trailing-stop exit. The model supplies a complete Backtrader strategy. QuantCode-Bench, the authors' earlier 400-task benchmark, then checks whether the file runs in a backtest, places at least one trade and passes a GPT-5.4 judge's assessment of whether it implements the request. Generations use greedy decoding. In agentic mode, a failed attempt receives structured feedback and the model can retry for up to 10 turns. T1 and T10 measure cumulative judge pass after one and ten turns.
The authors train on two kinds of material. Continued pretraining updates every parameter using raw code from GitHub repositories found through trading and backtesting keywords. The repositories were filtered on stars and recency, then MinHash-deduplicated. That left about 5 to 6 billion tokens, used for 2,250 steps with 512 sequences of 8,192 tokens. Supervised fine-tuning uses about 800 request-to-code pairs. Requests came from Reddit, StackExchange and GitHub and were ranked by a model-based judge. Material that did not read like a request was rewritten as a user query; benchmark tasks were excluded. An agent wrote the code targets with the framework in view, and targets survived only if they passed execution checks. Fine-tuning ran for three epochs at a learning rate of 1e-5. Qwen3.6-35B-A3B received both stages in sequence. Qwen3.5-397B-A17B received pretraining only.
The gain, and its limit
On the 35B model, single-turn backtest success rises from 46.0% at base to 53.8% after pretraining, then to 83.5% after fine-tuning. Judge pass follows the same path: 27.8%, 33.0% and 58.2%. For 397B, pretraining raises judge pass from 41.5% to 47.5% and backtest success from 65.5% to 71.5%.
Execution accounts for the largest jump. Fine-tuning adds +29.7 points at that gate. Before it, 46.2% of the pretrained model's strategies failed to run. After it, 16.5% failed to run, while a further 23.5 points ran without trading.
The authors applied fine-tuning only to the pretrained checkpoint, and only at 35B. They say the effect of fine-tuning without pretraining was not measured. Their account assigns the stages complementary roles against distinct failure modes, though this design cannot isolate their effects. The about 800 pairs therefore earn a conditional +25.2 points.
There are no confidence intervals or repeated runs. On 400 tasks, an unpaired standard error for a gap between two such rates is near 3.3 points. The pretraining gains of 5.2 and 6.0 points fall below two standard errors. Scores from the same tasks could support a tighter paired comparison, but the paper reports no paired test. The fine-tuning gain is too large for noise to matter. For scale, its 58.2% judge pass sits just under Gemini-3-Flash at 59.8%, and below GPT-5.4 at 70.2% and Claude Opus 4.6 at 75.8%.
What does a passing strategy establish?
The trade gate requires at least one trade. Fine-tuning lifts it from 36.8% to 60.0%, leaving 23.5 points of strategies that run without trading. The authors specify that "the objective is implementation correctness rather than profitability." They add that the success rates "must not be interpreted as evidence that the generated strategies are profitable or suitable for deployment." Those are sensible boundaries for a coding test. These figures say nothing about edge.
The report gives no period, instruments or bar frequency for the backtest data. Its semantic gate relies on an LLM, which the authors acknowledge may miss subtle logic mismatches. GPT-5.4 is also on the leaderboard at 70.2%; that overlap bears more on comparisons with frontier models than on comparisons within the Qwen lineage.
Repair, then a stalled agent
At T1, the 35B base passes 22.3%; by T10 it passes 47.5%. Pretraining raises the start to 32.3%, yet the checkpoint reaches only 32.5% after up to nine more turns. It finishes 15.0 points below base. Fine-tuning takes the model to 58.3% at T1 and 79.5% at T10.
The authors attribute the pretrained model's flat repair curve to weaker instruction following after training on raw repository code without chat-formatted data. They did not test that account with an ablation. We also could not resolve a discrepancy: the same base scores 27.8% single-turn and 22.3% at agentic T1. The paper offers no explanation.
Fine-tuning brings a repair gain of 21.2 points, in line with frontier models. Opus rises from 75.8% to 97.5%; GPT-5.4 goes from 70.2% to 95.0%.
Tool calls outside the strategy file
Both specialized 35B checkpoints often produced unclosed tool-call tags. The serving engine's parser could not extract a valid call, leaving neither checkpoint runnable on the authors' roughly 200-task repository track. A linear merge weighted 0.65 toward the fine-tuned checkpoint and 0.35 toward the original still gave malformed calls.
Recovery fine-tuning began with that merged model and used Qwen3.5-122B-A10B agent trajectories. Qualitative checks found well-formed calls again, though the authors retained no parse-success rate. With reasoning disabled, the recovered models solved 16.3% of repository tasks when trained on 16k-token sequences and 13.0% when trained on 32k. The base solved 23.7%. As the authors put it, "restoring syntax is not equivalent to restoring a tool-use policy."
The missing result for a developer plugin is a QuantCode-Bench score for either recovered checkpoint. The model whose strategies pass at 58.2% cannot drive tools; the models that can drive tools again have yet to be shown to retain that 58.2%. The paper has not reported both abilities in one checkpoint.
Our test stops elsewhere
We cannot reproduce the specialization. It requires large-scale continued pretraining, the validated training pairs and training compute we do not have. Our strategy uses an off-the-shelf model to generate code, so our evaluation tests none of the gains reported here.
The coding result from about 800 pairs over three epochs looks credible to us. A recovered checkpoint scoring close to 58% on QuantCode-Bench would change our view that the paper has yet to produce a deployable agent.