Paying a factor-mining agent to rewrite its rules bought no consistent portfolio IC in EverMine. Its Evolving arm finished statistically indistinguishable from an agent with a frozen rulebook, while the DeepSeek version used nearly three times the CPU. On Qwen3.8-27B, the gap was +0.74 thousandths of IC in development and -0.39 out of sample. Li, Zhang, Yao and their coauthors have isolated a useful question. Six Qwen runs and three DeepSeek-V4.1-Flash runs per arm cannot establish that editable rules have no value.
The pool keeps changing
The agent mines formulaic factors in the style of AlphaGen. It builds expression trees from OHLC, volume, turnover and VWAP, then submits them to an evaluator holding a 30-member pool. Each submission triggers a weight refit using L1-penalized least squares (lambda 0.005). Once the pool exceeds 30, the evaluator removes the member with the smallest absolute weight. It rejects a candidate outright when its signed correlation with any pool member exceeds 0.99. Portfolio IC is the time average of each bar's cross-sectional Pearson correlation between the combined signal and the next 10-minute close-to-close return.
The data cover Binance spot USDT pairs. Each month, the universe resets to the top 60 by trailing 30-day turnover. Development runs through 2025; the first half of 2026 is out of sample, scored with development weights frozen. A factor earns its place through its marginal contribution to the pool at submission. As that pool changes, a rule learned earlier may stop helping. Of 8,774 candidate evaluations across the 18 long trajectories, 8,215 (about 93.6%) occurred after the pool first filled.
The authors divide research state into Hist, the hypotheses, feedback and conversation; Frontier, the current pool; and Cap, editable documents containing skills, tools and procedures. Fixed retains full Hist while its initial Cap stays frozen. Evolving can edit Cap, provided it records where a rule applies and when it fails. Each trajectory faces matched caps, including USD 10 in model cost, 1,000 submissions and 16 CPU core-hours.
Did the edits pay?
Qwen development IC came in at 0.06164 for Evolving and 0.06090 for Fixed. Their +0.00074 gap has a 95% bootstrap interval of -0.00118 to +0.00268. Out of sample, the gap changes sign: -0.00039, with an interval of -0.00416 to +0.00318. DeepSeek's three runs per arm leave still more uncertainty. Its out-of-sample gap is -0.00164, with an interval of -0.01823 to +0.01309 against an IC level near 0.10. The authors caution that intervals at n=3 offer limited support for nominal coverage. The evidence supports no consistent gain from explicit capability evolution, which is also the abstract's claim.
The cost split is harder to dismiss. Across three runs per arm, DeepSeek Evolving averaged 16.0 CPU core-hours against Fixed's 5.8, and USD 4.30 against 2.34. It completed 642 evaluations to Fixed's 1,000. All three Evolving runs hit the CPU cap; all three Fixed runs hit the submission cap. Qwen's CPU spend was much closer, 7.23 versus 7.22 core-hours, though Evolving again completed fewer evaluations: 295.3 against 346.0.
Both models posted out-of-sample IC well above development IC despite frozen weights. Qwen, for example, was about 0.088 against 0.061. Its level rose by roughly 27 thousandths between periods, while the arms were separated by under one. The authors warn that the models may have encountered 2026 during pretraining.
What does the frozen rulebook lose?
The branch experiment gets closest to that question. At the 25% and 75% budget checkpoints in the six Qwen Evolving runs, the authors continue one branch with accumulated Cap and another with initial Cap. Both branches read Cap without editing it, and each has two repetitions, yielding 48 branches. At 25%, accumulated Cap added 0.00163 of IC versus 0.00203 for initial Cap. The difference was -0.00041, with an interval of -0.00110 to +0.00012; out of sample, it was -0.00126. At 75%, the development difference was +0.00004. By parent, development effects were positive twice and negative four times at each checkpoint. The two repetitions disagreed in sign in 5 of 12 cases.
Both branches could still read Hist. The test therefore measures what distilling experience into Cap adds when the agent already has that experience in context. It cannot settle the value of experience itself. Late branches changed little: across the five parents covered at 75%, final output correlated 0.9951 and 0.9924 with the starting pool.
A threshold follows the agent
Agents make up local screening rules. In one Qwen Evolving run, Batch A reused a residual-IC floor of 0.0015 that later entered Cap. The authors replayed two batches from the exact platform state. Of 12 candidates, 11 had been screened out, although 4 of those would have had positive marginal gain at the original pool. The accepted Batch A candidate added 0.00044635; the best Batch A reject would have added 0.00014081. Submitting every reject in sequence turned 3 of the 4 positives negative and reduced final IC by 0.0000175 in Batch A and 0.0000067 in Batch B.
"Independent value at the original state and realized value under sequential submission therefore answer different questions," they write. Yes, though this is 10^-5 of IC in one trajectory, analysed post hoc by the authors' own account. Their conclusion calls for separate scores for the priority experience selects, the alternatives it rejects and the realized portfolio outcome. Recorded experience should also be tied more directly to its actual use and to portfolio feedback.
Work on familiar formulas did produce gains. Parameter tuning was net-positive in 16 of 18 trajectories and accounted for 12.8% to 40.0% of each group's net improvement. At a 10^-4 gain threshold, it delivered 44 of 201 hits.
Repeated submissions show how much the pool matters. Among 178 exact-repeat pairs with a negative first gain, 50 became positive; 44 of those 50 were from Qwen Fixed. Of 311 pairs that began positive, 79 became negative. These counts cover only factors the agents chose to resubmit, as the paper notes.
Before trading the IC
The reported outcomes are IC. We found no signal-turnover, trading-cost or PnL figures, and the authors' related-work appendix notes that backtests depend on transaction costs. Features include the close at t; the target runs from that close to the next. A short-horizon reversal factor could pick up bid-ask bounce that a taker cannot collect. A tradable version would need signal turnover, long-short returns net of spread and impact, and Rank IC alongside Pearson, measured at matched evaluation counts.
We cannot reproduce the agent, its editable capability state or the branching framework. We are building a separate US equities baseline for fixed versus adaptive factor screening. Candidates enter either on a fixed standalone threshold or on their marginal contribution to the current pool. It tests portfolio-conditioned selection rather than the paper's agent comparison, and it has no results yet.
A replacement run would change my view if both branches lost access to Hist and accumulated Cap still beat initial Cap.