AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

FactorDGL's Mastercard edge shrinks under sequential evaluation

Jiao and Liu specify the optimizer, but graph causality and error units remain unclear.

2026-09-29 · 6 min read · Machine-learning return forecasting · US equities (MA and V)

Reviewing: FactorDGL: factor-driven dynamic graph learning for stock price forecasting · XueMei Jiao and ZiYang Liu · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

FactorDGL-Style Temporal Hypergraph Forecasting and Risk-Aware MA/V Selection

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (4)

  • Paper universe, June 2008–June 2024 Kaggle data and static train/validation/test dates.: Retain MA and V, but use available daily_prices and annual walk-forward evaluation from 2020-01-01 through 2024-07-01. (invalidates: Direct numerical comparison with paper Tables 2–5 and 9; Fig. 4 sequential errors)
  • Paper conducts forecasting only and specifies no execution, transaction costs, turnover or trading threshold.: Add validation-selected stock selection, next-day adjusted-open entry and adjusted-close exit, explicit per-fill costs and turnover reporting. (invalidates: Any attribution of net returns, turnover or trading profitability to the paper)
  • Paper leaves the initial residual, stage-network choice and α/β numerical weights unspecified.: Define the initial residual as normalized price minus learned observable-factor projection; select stage network and positive loss weights using past validation only. (invalidates: Exact replication of paper Tables 2–5 and 9; Fig. 4 sequential errors)
  • Paper observes share trading volume; requested input is daily_prices.volume_in_dollars.: Use observed dollar volume as the trading-activity prior factor, if selected on validation. (invalidates: Exact replication of paper Tables 2–5 and 9; Fig. 4 sequential errors)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 20.9%Jul 2024
Sharpe
0.70
Total Return
20.9%
Max Drawdown
-6.5%
CAGR
4.3%
Volatility
6.6%
Beta vs SPY
0.12
Trades
436

A 30% cut in yesterday's-price error on Mastercard would command a trader's attention. Jiao and Liu report 1.15 MAE for FactorDGL on the held-out set, against 1.65 for persistence. The harder comparison is with the best learned baseline: 1.23 for a temporal convolutional network (TCN) with attention. FactorDGL's 1.15 becomes 1.23 under the paper's sequential protocol, against 1.25 for TCN.

What FactorDGL does with a price series

FactorDGL makes each trading day in a single series a node. A node carries the min-max scaled price and "prior factors," including daily return and a moving average. Each hyperedge joins a day to its previous 30 days, roughly six trading weeks. Hypergraph convolution averages features within an edge and sends that information back to its member days. The propagation matrix is V = Dv^-1/2 H De^-1 H^T Dv^-1/2. Two layers with residual connections and LayerNorm yield an embedding for each day.

Cascading residual hidden-factor extraction (CRHFE) then works on the price left unexplained by the prior factors. At each stage, it projects the residual, joins that projection to the embedding, extracts a latent vector and subtracts a reconstructed portion of the residual. The procedure resembles boosting. Two stages run, and their outputs are concatenated. Two linear heads use the result: one forecasts next-day open, high, low and close; the other forecasts next-day 30-day rolling volatility and Sharpe, with a risk-free rate of zero. Training minimizes alpha times price MSE plus beta times risk MSE.

The authors' economic case rests on days moving prices jointly through volatility clustering and momentum accumulation, beyond the pairs considered by pairwise attention. Their data cover Mastercard and Visa daily bars from June 2008 to June 2024, with tests from January 2023 to June 2024. DJIA data run from 2000 to May 2019, with tests from 2017 to May 2019. On DJIA, FactorDGL reports 0.1705 MAE, versus 0.1983 for TCN and 0.2874 for persistence. The authors say "no trading strategy is implemented" and describe the results as "a focused validation rather than a universal proof of generalizability." Forecast superiority is the claim at issue.

The optimizer is specified

Adam uses learning rate 0.001, batch 32 and 50 epochs across five seeds. Validation searches hidden size {32, 64, 128} and learning rate {0.0005, 0.001, 0.005}. The window is L = 30; the network has two convolution layers and two residual stages. Training data alone set the min-max bounds. The first 29 rows are dropped instead of padded, and splits are specified to the month. The paper repeatedly takes care over leakage through the scaler and rolling indicators.

Can the hypergraph see future days?

Several implementation choices remain open. Alpha and beta are named twice without values. The stage network is "a small MLP or Transformer," with no size supplied. An "e.g." defines the initial residual, leaving the prior-factor map g unspecified. We did not find a stated window for the moving-average prior factor. The most consequential omission concerns causal propagation.

As written, H is T by T across all time steps, and V is symmetric. Day i belongs to hyperedges ending from i through i+30. A pass of V can therefore pool information from as far as 30 days after i. The authors say h_t uses "strictly historical information," yet we did not find where they restrict the graph to a prefix. A causal-prefix graph built for each forecast would address that issue. It would also differ from the model described by the equations and might fail to reproduce the table.

The errors do not reconcile

Every MAE and RMSE is said to be "calculated on the normalized price values." If the bounds are fixed on 2008-2020, an MAE of 1.15 puts the typical one-day error beyond the entire training price range. The units appear to have been misreported; the tables give readers no way to recover dollars.

Persistence is deterministic, yet its Mastercard result has ±0.05 attached. ARIMA's grid includes the random walk, but its R² falls below persistence: 0.85 against 0.86 on Mastercard, and 0.7895 against 0.8240 on DJIA. Another conflict appears in the ablation table. Its optimal configuration, L = 30 with two layers and two stages, records 1.23 MAE and 1.75 RMSE. The full model, described in the text as the same configuration, records 1.15 and 1.68. The receptive-field table also gives HTM at RF = 30 a result of 1.15. Those entries cannot all represent one model.

The protocol difference matters even more. On Mastercard, the authors attribute a shift from 1.15 to 1.23 to "different evaluation protocols." They caution that held-out MAEs "should not be directly compared" with sequential MAEs and identify the sequential setting, with its smaller edge over TCN, as more realistic.

We compare the two with that warning in view. The held-out setup already describes test predictions as "generated sequentially based solely on past information." The sequential evaluation uses nearly the same description. "different evaluation protocols" leaves the operative difference unclear. If both runs are causal, the paper's account does not explain the 0.08 gap, equal to FactorDGL's entire held-out Mastercard edge over TCN.

The risk forecasts also lack a telling comparison. Next-day 30-day volatility shares 29 of its 30 returns with today's measure. Its reported R² is 0.940, against 0.920 for TCN, but overlap alone could account for much of that fit. Without persistence, the table cannot show how much. Tables 4 and 5 have no persistence row and do not identify their dataset.

A sliding window over one stock

The dynamic graph in the title is a fixed-shape, sliding 30-day window on a single series. It connects no other stock. The substitution study gives the hypergraph operator a fairer comparison: receptive fields are matched, and the hypergraph module (HTM) has 106.4K parameters versus 204.9K for local attention.

For Mastercard at RF = 30, HTM's 1.15 beats local attention's 1.28. At RF = 10 and RF = 50, the respective results are 1.29 against 1.33 and 1.27 against 1.30. Nearly all of the Mastercard advantage appears at the headline window. The gap stays roughly constant across windows for Visa and DJIA. Visa records 1.39 against 1.41, 1.35 against 1.38 and 1.37 against 1.40 at RF = 10, 30 and 50. DJIA records 0.1786 against 0.1865, 0.1705 against 0.1768 and 0.1738 against 0.1819.

Without HTM, the result is 2.05, worse than persistence at 1.65. The remaining stack cannot match a repeated price on its own, despite taking today's price as an input. Repeating that price scores 1.65.

Our Mastercard/Visa trade

The paper reports no trading returns to compare with ours. Conversely, our run did not calculate its forecast MAEs. We added a trading rule: after each close, the model forecasts next-day close and volatility for MA and V, then divides forecast return by forecast volatility. We buy the highest-scoring stock if its score clears a threshold selected on the prior year's validation data; otherwise, we hold cash. Position size cannot exceed 50% of initial capital. We refit annually on expanding history, starting with training through 2018 and validation in 2019. Our implementation choices included stock-specific causal-prefix graphs, validation-chosen loss weights and stage network, our own residual initialization, and dollar rather than share volume.

From 2020-01-01 to 2024-07-01, the backtest returned 20.91% net of commissions of $0.004 a share, subject to a $1 minimum. At most half the capital was invested. Annualized volatility was 6.62%, beta to SPY was 0.12, Sharpe was 0.70 and maximum drawdown was -6.46%. The return is modest in light of that limited deployment.

The fill assumption is the biggest caveat. Our rule enters at the next open and exits at that day's close, while fills in the daily-bar run were recorded at the close. The 20.91% may therefore differ from the return on the open-to-close trade we designed. We did not model market impact, borrow or financing. This was one automated pass over a two-name universe. It reflects our choices more than Jiao and Liu's architecture and can neither confirm nor refute their 1.15.

Released code that constructs a prefix-restricted graph and still beats 1.25 under the sequential protocol would change our view. For now, FactorDGL reads as a TCN-class forecaster: its 0.08 held-out Mastercard edge is no bigger than the unexplained difference between its own protocols.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each annual refit, fit stock-specific scalers and the joint price/risk model on eligible expanding-history samples; select model settings and a score threshold using the preceding validation year only.
After close t, forecast t+1 adjusted OHLC, trailing volatility and trailing Sharpe for MA and V.
For each stock, score = (forecast close[t+1] / observed close[t] - 1) / forecast volatility[t+1].
Keep stocks with positive forecast return, positive finite forecast volatility and score above the frozen threshold.
Buy the highest-scoring eligible stock, or hold cash; cap the position at 50% of initial capital and check the 4× leverage ceiling.
The specification calls for entry at the next adjusted open and exit at that day's adjusted close, subject to available actual fills. The platform records daily-bar MOC execution instead; reconcile the executed blotter before interpreting returns as open-to-close trades.