A 127.17% human record says little about an agent stack that was never built. Fang's paper presents a design with one idea worth keeping, yet supplies no evidence for the system itself. The abstract admits the gap: "That record is evidence for the underlying method rather than for any AI system; no implementation is evaluated here." Its defence of the architecture carries the argument. The design "is not a speculative design: it encodes a method the author first executed by hand." The question is whether hand execution validates the proposed encoding.
We ran nothing against this paper, and no figure below comes from any test of ours. Reproduction is impossible under the stated framework because Fang deliberately withholds the probability calibrations, scoring weights and fusion parameters. Required inputs extend beyond price, filings, news and macro data to expert scientific and regulatory judgment. An honest market substitute is unavailable as well. A U.S.-listed biotechnology version would remove the mechanism for reconciling valuation, regulatory velocity, liquidity and pricing across China, Hong Kong and the U.S. The paper calls that mechanism its second stated contribution. The discussion below therefore rests entirely on the paper.
What Fang has designed
Fang identifies a real category error. Every multi-agent investment system reviewed in the paper, TradingAgents, ai-hedge-fund, AlphaAgents and FinRobot, values companies through earnings and cash flows using DCF, owner earnings, P/E or EV/EBITDA. Those inputs do not exist for pre-revenue clinical-stage biotech. Value instead depends on binary gates. The cited base rates show their importance: fewer than 14% of drugs entering Phase I are ultimately approved, while roughly half of new molecular entities fail approval on first submission. A DCF agent can still produce a precise figure, though its method cannot generate the underlying input.
The remedy is a layered agent stack. Scientific and clinical agents assess mechanism, trial design, endpoints and regulatory path, then issue structured views with explicit confidence levels. A valuation agent converts those views into an rNPV, meaning a net present value that multiplies each stage's cash flows by the probability of reaching that stage. Its output is a range. Cross-market agents, one each for A-share, Hong Kong and U.S., translate the same asset into each venue's conventions while leaving the pricing spread exposed. A risk agent sizes binary positions and estimates correlation across shared mechanisms and regulators. The synthesizer agent combines the views. Conflicts caused by an unexamined issue go back for re-analysis.
Beneath the stack is Fang's "Glocal" practice, a three-dimensional discretionary method. Dimension 1 reassesses founder capability and global operating resources against Fang's chosen peer set. Dimension 2 targets a cross-border difference in regulatory and clinical speed by shortening assumed time-to-value, then sizing above consensus. Dimension 3 limits a single holding to roughly 10% of the portfolio and maintains roughly a 20% liquidity reserve.
Figures appear for only one fund. Fang calls it China's first dedicated cross-border biotechnology fund, CSRC code 001984, and served as sole PM from its February 2019 inception. The launch followed HKEX Chapter 18A, effective on 30 April 2018, by about ten months. Within sixteen months, the fund returned 127.17% against a 50.67% benchmark. Its share of the platform's cross-border AUM grew 93-fold. It also ranked first of 276 comparable peer funds through the 2020 stress period, when the peer average return became negative.
The paper gives no backtest, dataset or universe. Its synthetic "Company X" Phase II walkthrough is expressly excluded as "a performance claim, a backtest, or a description of any actual company, position, or decision."
The practice has a record. The encoding does not
Taken at face value, the performance supports a human process used in one vehicle over sixteen months. Fang then claims that agents can decompose and reproduce that process. With nothing implemented, the paper offers no evidence for this step.
The element credited with the standout result requires no language model. Fang attributes the 1/276 ranking through 2020 to Dimension 3's liquidity discipline, specifically a 10% single-name cap and a 20% reserve. A risk system can enforce both constraints.
No scientific analyst agent is needed for that job.
Return attribution is even less developed. The sixteen months beginning in February 2019 fall inside the opportunity created by Chapter 18A, which Fang says "opened an entire asset class to public capital overnight while it remained largely unpriced by the market." The paper addresses this timing problem twice. First comes the launch date. The vehicle began "preceding the asset class's formal benchmark index by nearly a year, and proving that the fund captured, rather than followed, market recognition." A launch ten months after the reform and a year before the index shows that the fund arrived early. It leaves stock selection mixed together with sector-wide repricing, and the return has no accompanying risk-adjusted figure.
The second response is the paper's sole attempt at separation. Applied concurrently from August 2019 to a "structurally distinct domestic-market fund," the same method allegedly "produced comparable top-tier outperformance." No return, benchmark or period length accompanies the statement. This repeatability claim carries most of the answer to favorable timing, yet has no figures behind it.
The index responsible for 50.67% goes unnamed. Fang also leaves unclear whether 127.17% is gross or net of fees. We found no Sharpe, drawdown, volatility or beta in the text. The 1/276 rank captures one cross-sectional snapshot, and the peer distribution has no reported dispersion.
Proprietary inputs leave a workflow
Fang openly withholds proprietary probability calibrations and weighting parameters for professional confidentiality. That may suit the business, but readers receive a workflow rather than an investment rule. The cited rNPV literature says milestone probability dominates the valuation, while base rates differ by more than an order of magnitude across indications and phases.
The paper also acknowledges where the calculation comes from. "The underlying arithmetic follows the risk-adjusted net present value and real-option traditions developed for pharmaceutical assets," while "the framework's contribution is not this arithmetic but the agentic pipeline that supplies its scientific inputs." Once calibration is withheld, the valuation layer becomes an arithmetic template that is already available. The proposed contribution lies in the agent chain, whose calibrations and fusion weights are withheld as well. Disclosure reaches only the agent roster and execution order.
Keep the conflict taxonomy
One proposition survives these weaknesses: "the architecture must preserve disagreement, rather than average it away." Fang's conflict typology deserves to be used. When a strong mechanism assessment meets a weak approval-path assessment, the conflict belongs in the probability estimate. When strong efficacy data meets a stretched valuation, it belongs in the margin of safety. Combining both disputes into one average score discards information a PM needs. The point stands with or without LLMs.
Fang names three future directions, and only the first amounts to a test: validate valuations against realized outcomes on disclosed events. That work would matter. The agents' milestone probabilities should be calibrated against realized read-outs, followed by incremental performance against a generic multi-agent baseline, net of costs. My threshold is a calibration curve covering 200 disclosed Phase II and Phase III outcomes. The paper reports nothing comparable. Evidence at that level would change my view of the entire stack.