An 11.14% live gain while CSI300 lost 6.10% could get a strategy funded. Liu and Chen give good reasons to hold off.

The weekly book

QuantPits ranks a subset of the CSI300 under its own eligibility rules. For each roughly weekly batch, an ensemble scores eligible stocks using data through the preceding close. The paper withholds the model identities. A commit dated 3 June 2026 did, however, make percentile-rank normalization the default way to combine their scores. Each batch produces three tiers: primary picks, alternates and a candidate list about three times the primary count.

The TopK Dropout book holds about 22 names. Under the configured Drop3 rule, each batch removes low-ranked holdings and fills the openings from the list.

Chen led manual execution. At the open, designated positions were sold in full; cash was divided roughly equally among buys in 100-share lots. The operator judged suspensions and limit states by hand, substituting alternates when a pick could not be bought. Friction depended on the side of the trade. For the year, notional-weighted Sell friction was -0.3839%, compared with +0.0348% on Buys, before explicit fees of about 0.0528% of notional. Negative means adverse. The operator also says LLM agents wrote most engine changes in response to human-issued tasks. There is no authorship count to support that account.

How much does the year prove?

Across 241 daily observations from October 2025 to September 2026, the account gained +11.1391% net of recorded settlement costs. CSI300 returned -6.0997%, and Sharpe was 0.6726. Single-factor annualized alpha reached +14.4911%, with a HAC(3) p of 0.27526. Every covariance choice the authors try puts it between 0.27526 and 0.32239. With liquidity, momentum and volatility proxies added, the intercept is 13.52% at p=0.226. The authors acknowledge the null result in their abstract and do not argue around it.

They divide the year after the fact into three windows of 94, 82 and 65 days. March to June lost 14.7887% while CSI300 gained 5.7058%. Market-only alpha for that window was -58.1058% (p=0.01778). With the three lagged style proxies, it becomes -25.3155% at p=0.22919; the 95% interval stretches from -66.9051% to +16.2741%. The authors regard that shift as compatible with exposure explanations without naming a mechanism. The account still lost 14.7887%. July to September brought a 15.2461% gain as the index fell 12.4876%.

HAC does not fix the choice of windows, as the authors say plainly. They computed no Deflated Sharpe Ratio or PBO. Even the early window changes its statistical label with the covariance choice: its +30.07% alpha clears 5% under HAC(3) (p=0.03339), yet misses under OLS (p=0.09396).

June against random rankings

QuantPits lost 9.2297% in June 2026 while CSI300 rose 1.7847%. The authors compare it with a Drop3 random-ranking reference: K=22, 5,000 annual paths and constituent updates every five trading days. June's median return in those paths is -5.1439%. Only 272 paths (5.44%) did as badly as production or worse. For the full year, 51 of 5,000 paths (1.02%) matched or exceeded production's +11.1391%.

A bottom-tail month and a top-tail year.

The reference ignores lots, minimum fees and opening fills; it can also buy suspended names. Its eligible subset had a gross equal-weight return of -5.1523% for the year. The authors warn that the tail fraction "is not a competing calibrated alpha test, skill probability, or causal attribution test."

Those checks help pin down what happened in June. May entries cost 2.84050 points, against 5.06181 from stocks outside the 29 May cohort, weakening the story that May entries alone drove the loss. Trading less offers no simple answer either. In the March-to-June replay, Drop0 lost 16.6910%, Drop1 19.6462%, Drop3 12.9816%, Drop6 11.8967% and Drop9 17.0115%. A basket frozen at the 2 March close lost 16.6910% against production's 15.4779%, making the freeze 1.2131 points worse. Freeze at the 1 April close instead, and the basket loses 7.2509% against production's 12.0480%. The authors say later decisions could have hurt, without identifying a responsible component.

Which code produced the orders?

The inspected engineering records contain seven dated entries from 10 March through 28 June. Four March commits altered ensemble alignment and normalization, the prediction loader, predict-only artifacts and per-mode recorder lookup. A 27 May note describes hyperparameter revisions and retraining. The fusion default changes on 3 June; a 28 June note documents a new CPCV training capability. The runtime inventory contains 50 run records (49 declared successful, one failed) and six cycle records, only one sealed complete. No complete recommendation-ID-to-producing-run link survives.

The authors give a hypothetical order whose remaining records fit both March loader designs. They have not checked whether any actual order has that problem. To describe the gap, they use compatible histories, meaning deployment timelines that all fit the surviving records. A claim is archive-identified only when it has the same value in every such timeline. Recovering a run record for their example would identify the loader that produced the order. It would leave open whether the other loader could have prevented a later portfolio loss. Provenance "is not causal magic."

A 26 June LLM Critic trace proposed changing fit_end_time. Under the recorded slide mode, that field did not control the training boundary. A note two days later says the proposal was skipped.

Recording the missing links

The authors propose storing each recommendation ID with its run ID, then tying that run to immutable code, configuration, artifact and input snapshots. Generation time and an output digest would travel with the record. Fills would link separately to a recommendation, a substitution or an explicit manual status. The authors call the proposal unvalidated; this year's missing bindings prompted it.

We could not test any of this ourselves. The universe is Chinese equities outside our US price coverage. The authors withhold trade-level records, and the archive never had complete run-to-order links. Reproducing the random reference also requires the unpublished eligibility rules.

From a 28 June 2024 baseline, the authors' passive cash, index, bond and gold pool made 30.64%, against 27.75% for the quant account. They call that informal context. I would read a prospective year as evidence of skill if it were logged under their protocol and its alpha cleared the bar that this year's p=0.275 missed.