The execution rules are worth borrowing. The 2.08 Sharpe is harder to trust when the paper gives two returns for the same hedged book.

What gets traded

Zeng, Ding and Yi trade the spread between Shanghai and New York gold. SHFE AU trades in CNY per gram; COMEX GC trades in USD per troy ounce. They use USD/CNH and 31.1035 grams per ounce to put GC in CNY per gram, then take the difference as their adjusted spread. A trade needs both a stretched spread z-score and agreement from a neural forecast. The book buys the cheap market and sells the rich one, seeking convergence.

That pair carries currency exposure. Long SHFE and short GC means long gold in CNY and short gold in USD. The authors hedge with SGX USD/CNH futures (USD 100,000 notional), rounding the COMEX notional to whole contracts. They report the remaining CNH/CNY basis and rounding error as a separate PnL line.

Their LTA forecaster combines a bidirectional LSTM, a Transformer encoder and attention pooling to predict the next-day direction of the USD-adjusted gold return. The daily data run from 2018 to 2024. Calendar splits assign 2018-2021 to training, 2022 to threshold setting, 2023 to model comparison and 2024 to frozen-model paper trading.

Execution starts at the next session, with one tick of slippage on each leg. Holidays and limit moves are skipped. If one of the three legs fills and another fails, the filled leg is flattened.

The paper calls its headline walk-forward table a backtest/forward-paper result. It reports 14.8% return, 7.1% vol, Sharpe 2.08 and 4.2% max drawdown for the hedged book, across 52 trades. The table has no date, though its 52 trades match the sum of the 2024 monthly counts (4+4+5+3+5+4+4+3+5+5+4+6). Unhedged, the book returns 11.2% on 9.0% vol, with Sharpe 1.24 and 12.7% drawdown.

Execution is the useful contribution

Cross-border spread profits depend on fills and costs, and the paper gives them unusual attention. Contracts roll 5 business days before expiry. Settlement prices generate signals, never assumed fills. The basis residual has its own -0.4% line. The failed-leg rule has a large effect: removing it takes hedged return from 14.8% to 10.6% and raises max drawdown from 4.2% to 7.4%. A daily cross-exchange spread trader could use these rules directly.

How much does LTA add?

LTA gets 63.7% accuracy (±1.2 across 10 seeds) and MCC 0.274 on 237 labelled days in 2023. The plain Transformer gets 61.4% (±1.5) and MCC 0.228. The authors acknowledge that the 2.3-point difference is not significant at 5%: "the evidence does not prove a statistically dominant architecture at the 5% level". They still describe the hybrid as "useful for timing and position filtering". We did not find McNemar or Diebold-Mariano statistics in the paper. Its ablation is described in words, without figures.

The PnL attribution gives forecast timing +5.4%, the paper's only quantified measure of that usefulness. This exceeds half the +10.1% assigned to spread convergence, yet the paper leaves the +5.4% calculation unexplained. Given the insignificant 2.3-point accuracy edge (63.7% vs 61.4%), a Transformer filter might earn the same credit. We did not find a run of the z-score rule with the model switched off, the test that would separate the forecast's contribution.

The returns disagree

One performance table says 14.8% for the hedged book. The attribution totals +15.3% and calls that the forward paper-trading total. By our arithmetic, compounding the twelve monthly returns also gives about 15.3%; those months range from -0.6% in April to +3.4% in December. The paper leaves the monthly table unlabelled as to hedged or unhedged, although its compound return matches the hedged attribution. The figure that differs is the one paired with the Sharpe.

Costs raise another question. The attribution deducts -1.5% for costs and rolls. Multiplying baseline costs of 1.5% by 2.5 suggests roughly another 2.25 points of cost, yet the 2.5x scenario loses 6.7 points of return, ending at 8.1%. A move to 2-tick slippage cuts return 5.4 points to 9.4%, more than the whole baseline cost line.

The monthly returns look smooth. We calculate standard deviation of about 1% a month, or about 3.6% annualized, versus the stated 7.1% vol. August's -3.8% worst drawdown suggests intra-month movement could account for some of the difference. The win rates also resist reconciliation with 52 trades: thirty-two wins imply 61.5%, and 35 imply 67.3%, against reported rates of 61.3% and 67.8%. Both Sharpes equal return divided by vol exactly (14.8/7.1, 11.2/9.0); neither calculation deducts a risk-free rate.

The hedge adds 3.6 points of return over the unhedged book, while its FX hedge attribution is only +1.7%. The authors cite adverse currency moves for the unhedged book and a second effect, "when reduced volatility allows the same risk budget to be deployed more consistently". Trade count stays at 52. Volatility-targeted sizing, mentioned in the robustness section, is therefore the most plausible source of the other 1.9 points; the attribution gives sizing no line of its own. The hedge case chiefly rests on lower volatility (9.0% to 7.1%) and drawdown (12.7% to 4.2%). Its unexplained return still enters the 2.08 Sharpe.

These discrepancies matter together because the paper calls its evidence "reproducible". The authors also write that the complete ledger "should be supplied as supplementary material before the results are treated as fully reproducible". The abstract says the result is "not proof of durable arbitrage profitability without intraday executable quotes and complete trade-level records", while calling it "preliminary, reproducible evidence of economic usefulness". The abstract and conclusion's reproducibility claim is difficult to square with returns that differ across tables. The full ledger is available on request. Its five displayed rows are marked "example format", leaving readers unable to tell whether T14's -0.41% records an actual trade.

The authors acknowledge that 52 trades cannot establish long-run robustness.

Why we could not rerun the spread

We have COMEX GC, but lack SHFE AU, SGX USD/CNH futures and the CNH/CNY basis series. The Shanghai leg is indispensable to this spread. GC or a US gold ETF alone would leave an outright gold position, a different trade. US minute bars cannot establish fills across Shanghai, New York and Singapore sessions either.

The ledger would change this assessment

The authors specify what the full ledger should contain: signal time, contract months per leg, entry and exit, fees, margin currency. Publishing those 52 rows against one reconciled annual return, along with their proposed five-minute quote check on signal days, would answer most of the questions here.

Until then, use the execution protocol and set the 2.08 aside.