AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

MUFASA's 0.4465 S&P 500 error has no portfolio test

Its valuation formulas beat 17 baselines on MAPE, while the role of current-price multiples remains unclear.

2026-09-29 · 7 min read · Fundamental Value · US listed equities

Reviewing: Self-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental Analysis · Kelvin J. L. Koa, Filip Orestav, Shengqiong Wu et al. · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Walk-Forward Sector- and Regime-Weighted Fundamental Valuation

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

Run on a different market than the paper

Restrict the paper's US, UK, and Chinese equity universe to US equities. Accounting-based valuation and mispricing signals remain applicable, but the resulting test cannot establish the reported cross-country generalization.

The paper's own figures describe its universe and do not carry over to ours.

This is not a replication of the paper

  • The described data and backtest do not provide the paper's LLM-agent orchestration and semantic-retrieval system. A reproducible implementation could use classical symbolic regression with fitted sector- and regime-dependent weights; its results would test that substitute, not MUFASA's multi-agent discovery claim.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (6)

  • Table 4 S&P 500 earnings agent: NI_ps · (1 + 0.01 · g_EPS) · PE · 0.88: Preserved verbatim for retrospective reproduction only; live trading discovers and refits an earnings expression using historically available data and peer, rather than own-stock, multiples. (invalidates: Table 2 S&P 500 MUFASA MAPE 0.4465 as an expected result for this test; Table 3 S&P 500 full-MUFASA MAPE 0.4465)
  • Table 4 S&P 500 cash-flow agent: FCF_ps · (EV/EBITDA) · (100/88.38): Preserved verbatim for retrospective reproduction only; live cash-flow expression and coefficients are fitted on historically matured observations using available peer multiples. (invalidates: Table 2 S&P 500 MUFASA MAPE 0.4465 as an expected result for this test; Table 3 S&P 500 full-MUFASA MAPE 0.4465)
  • Table 4 S&P 500 growth agent: (P/FCF)_mkt · (FCF_ps + 0.29 · (P/FCF)/(P/FCF)_mkt) · (1 − 14.80 · (FCF_ps < 1.08)) · 0.01: Preserved verbatim, including its indicator, for retrospective reproduction only; live growth discovery refits structure and coefficients and excludes the forecast stock's own price-derived P/FCF. (invalidates: Table 2 S&P 500 MUFASA MAPE 0.4465 as an expected result for this test; Table 3 S&P 500 full-MUFASA MAPE 0.4465)
  • Paper forecasting resolution: quarterly, with the stock price at the end of the fiscal quarter immediately following each filing period as target; first 80% of quarters train and remaining 20% are held out.: Retain the eligible quarterly target, but make daily trading decisions and use chronological, expanding, matured-target-only walk-forward fitting instead of one fixed 80/20 split. (invalidates: Table 2 held-out MAPEs for all five paper markets; Table 3 S&P 500 ablation MAPEs; Appendix G Figure 7 aggregate error percentiles)

2 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 37.0%Jul 2024
Sharpe
0.41
Total Return
37.0%
Max Drawdown
-44.3%
CAGR
7.3%
Volatility
20.3%
Beta vs SPY
0.66
Trades
15,438

What the paper reports for its own strategy

  • Out-of-sample (last 20% of quarters) next-quarter price-level MAPE, not trading performance; no portfolio, Sharpe, returns or cost assumptions are reported. - S&P 500 (1990 Q1-2023 Q2 sample): 0.4465 - Russell 2000 (1990-2025): 0.5718 - STOXX 600 (1995-2025): 0.5064 - FTSE APAC ex-Japan (1996-2025): 0.5814 - CSI 300 (2002-2021): 0.4343 - All-data average: 0.5081

A better next-quarter price forecast gives a trader no position to hold. Koa, Orestav, Wu, Wooldridge and Huang report forecasting accuracy, measured by mean absolute percentage error (MAPE), and claim no trading edge. Yet their introduction aims for interpretable formulas linking financial variables to future returns, while the fitted target is a price level. MUFASA reaches a test MAPE of 0.4465 on the S&P 500, best among seventeen baselines. The paper reports no portfolio, Sharpe, turnover or cost assumption. For that price-level result to matter to stock selection, it needs a portfolio test and a last-price benchmark. Both are missing.

How the five agents build a price

The premise allows more than one sensible valuation equation. An earnings investor and a book-value investor may value the same stock differently, with their relative success changing by sector and market regime. MUFASA (Multi-Agent Fundamental Analysis with Symbolic Adaptive learning) gives five LLM agents on Qwen3-4B-Thinking separate remits: earnings, cash flow, asset, growth or quality. Each proposes a symbolic valuation formula, retrieves Bloomberg variables, runs the formula as code and fits its coefficients by differential evolution. The score is MAPE against the stock price at the end of the following fiscal quarter.

Each agent remembers earlier equations and their statistical summaries: p90 and p95 error, coefficient of variation, Spearman rank correlation and signed error by price quartile. It turns that history into text lessons. A meta-coordinator agent uses SLSQP to weight the five outputs by sector and bull/bear context. A 20% rise from a trough declares a bull market; a 20% fall from a peak declares a bear. The weighted sum is the final valuation.

The authors use original quarterly filings, without restatements, across five universes:

For each universe, the first 80% of quarters train the system and the last 20% are held out. The implied trade, buying stocks whose estimated value exceeds their price, remains unexecuted.

The win over 17 baselines

MUFASA averages 0.5081 MAPE across all data. PySR records 0.6351, Random Forest 0.6579 and the sector-median P/E 0.6652. MUFASA wins in every market, with an average reduction of about 20% against PySR. We could not reconstruct the paper's stated 19.65% from its results table. Some per-market margins are narrow: on the S&P 500, LLM-SR scores 0.4761 against MUFASA's 0.4465, a gap of about 6%. On the Russell 2000, FinVision's 0.5842 trails MUFASA's 0.5718 by about 2%.

The comparisons tell a less tidy story than the headline. Classical genetic-programming SR, PySR, averages 0.6351; the LLM versions score 0.7271 for LLM-SR and 0.7991 for FunSearch. The authors suggest scientific priors transfer poorly to finance, a plausible reading. Their ablations make the case for separate perspectives more directly. A single unified agent scores 0.4971 on the S&P 500 and 0.8994 on the Russell 2000. Equal weights across specialists bring those figures to 0.4652 and 0.8768, improving results in four of five markets. Selecting the best single specialist reaches 1.9173 on the Russell 2000.

Statistical memory, the third of the three components named in the abstract, has a patchier showing. With the meta-coordinator present, plain binary pass/fail memory beats MAPE-only memory on the S&P 500 (0.4646 vs 0.4702) and Russell 2000 (0.6445 vs 0.7525). Full statistical memory wins overall. On the two US indices, though, moving from binary to scalar feedback worsens the result.

There are discrepancies in the reported figures. The single-perspective Russell result is 0.8994 in the main ablation table and 0.8987 in the extended table; equal weights are 0.8768 in one and 0.8831 in the other. For the S&P 500 comparison with PySR, the Diebold-Mariano table gives a difference of -0.2695, while the MAPE table implies -0.1929. The DM tests cover financial LLMs, valuation ratios and cash-flow models, and SR models. They omit Random Forest, the second-best baseline on average, along with both agentic baselines. The authors conclude that random variation mostly does not drive the gains. Their DM table leaves the closest baseline untested in two markets: FinVision on the Russell 2000 (0.5842 vs 0.5718) and FinCon on FTSE APAC (0.6358 vs 0.5814).

Whose P/E enters the equation?

The answer matters to a trader. On the S&P 500, the learned earnings equation is NI per share × (1 + 0.01 × EPS growth) × PE × 0.88. The Russell 2000 asset equation is TBVPS × PB × (1 - 0.24 × (PB - 1.00)); it multiplies tangible book value per share by PB. If PE and PB are each firm's current ratios, the formulas largely reconstruct today's price. The resulting forecast would be close to 0.88 times the current quote, leaving little information in a value-over-price gap.

The variable definitions leave the source of those multiples unresolved. Distilled agent learnings identify SECTOR_MEDIAN_PE and SECTOR_MEDIAN_P_BOOK as anchors, suggesting peer multiples. With peer multiples, the formula becomes a relative-value forecast. The observed errors also weigh against pure price reconstruction: Fin-R1's EPS_TTM × PE scores 0.5263 on the S&P 500. A formula reproducing today's price would be unlikely to miss the price one quarter later by half on average.

Either reading leaves a gap.

We did not find a last-observed-price benchmark among the seventeen baselines. A reader therefore cannot judge 0.4465 against the obvious one-quarter comparison: the last observed price.

Timing and the 27-quarter test

The target is the price at the end of the fiscal quarter following each filing period. A filing becomes public some time after the quarter it describes closes. We did not find a release-date lag in the setup, so the paper's information set differs from a daily investor's. Using original rather than restated filings does remove look-ahead from later restatements.

Regime dating poses another timing question. The wording says a bull market is declared once the index is 20% above its trough, which reads like a real-time crossing rule. Pagan and Sossounov dating, which the paper cites, assigns trough-to-peak phases after the fact. The paper does not specify which approach it computes. Only the crossing can be observed in real time, and the meta-coordinator conditions its weights on the regime label.

The held-out window is short. By our arithmetic, Twenty percent of the S&P 500's 134 quarters comes to about 27 quarters. The authors acknowledge the small number of test quarters per market and apply the HLN correction. They report no run-to-run variance for a 4B model sampling at temperature 0.7. Some equations also strain the interpretability claim: the S&P 500 asset formula assigns R&D/Assets a coefficient of 13829.10, while the STOXX growth formula raises revenue per share divided by 809734.00 to the power 34086.20.

Our US-only substitute

We could not run the paper's system: we had no LLM-agent orchestration or semantic retrieval layer. We restricted the universe to US equities as well, so our figures say nothing about its cross-country results. Our adaptation uses deterministic candidate search in place of the Qwen agents, with a separate search for each perspective (earnings, cash flow, asset, growth, quality). We fit coefficients by differential evolution and set nonnegative SLSQP weights by sector and S&P 500 bull/bear regime. Regimes switch at the 20% crossing. Each walk-forward refit uses only targets that have already matured.

Our universe contains 300 capitalization-ranked US stocks, observed daily from 2020-01-01 to 2024-07-01. Each day, we rank them by forecast value divided by adjusted close, minus one. The book holds the top decile with a positive gap, long only and equal weight. Positions have a 10% cap; orders are limited to 1% of daily dollar volume. We charged $0.004 a share, subject to a $1 minimum per order. Fills occur at the closing auction, although we intended next-open fills. We have not reconciled those two execution assumptions.

The figures printed above this text are ours. The paper supplies no Sharpe, return or drawdown to compare with them. Its 0.4465 is a price-level MAPE from a different method and setup; our run produced no forecast MAPE, so we cannot tell whether the substitute matched its accuracy.

Read the risk alongside the gain. From 2020 to mid-2024, our book returned 36.96% with 20.29% annualized volatility and a worst drawdown of 44.26%. This is one automated pass through a deterministic search and our execution choices, rather than a verdict on the authors' work. It says nothing about MUFASA's multi-agent claim.

Would the decile survive costs?

The authors have the inputs for a stock-selection test on their own S&P 500 test quarters. They could sort each quarter by MUFASA value over price, using a price available once the filing is public, and report next-quarter decile returns. A last-price row belongs beside the MAPE results. A positive decile spread that survives costs, with sector-median multiples in the formulas, would move us. A MAPE that fails to beat last price would move us the other way. Until those comparisons exist, 0.4465 remains an accuracy result untested as a trading signal.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each decision close:
  Determine the SPX bull/bear regime from observed index closes; abstain until declared.
  Use identifiable original filings available by the decision, complete published TTM histories,
    split-aligned shares and leave-one-out peer multiples.
  Refit perspective-constrained specialists on targets observable by the fit cutoff;
    update statistical memory and select specialists by sector/regime.
  Fit nonnegative context weights summing to one; abstain where weights are not estimable.
  Rank eligible stocks by combined forecast / observed adjusted close - 1.
  Target equal-weight positions in the highest decile with positive gaps, subject to
    a 10% position cap and 1% of observed daily dollar-volume order cap; exit others.

The specification calls for next-eligible-open execution, but the supplied cost record identifies the reported backtest's execution as daily-bar MOC. The blotter must be checked before attributing these results to next-open fills.