A daily lag choice beats every heavier trading model the authors test on nine sector ETFs. Their plain vector autoregression is the trade worth examining.
What Tan and coauthors built
The premise is familiar: information can reach one sector before another. Yesterday's financials or energy return may help forecast today's industrials, though the useful lag can shift with the regime. VIX and other macro states can also move several sectors together, making a shared driver look like a lead-lag link.
Tan, Zanoria, Chen, Lyu and Cucuringu seek the causal lag structure: which series drives which, and when. Their ORACLE-VARX procedure has three steps. Double/debiased machine learning (DML) first removes effects attributed to lagged confounders. One of five learners predicts each series and each of its lags from those confounders: Extra Trees and Random Forest are tree ensembles; LightGBM and XGBoost use gradient boosting; TabPFN is a pretrained tabular transformer. OLS then estimates the VAR coefficients from the residuals. Next comes ACLE (adaptive causal lag estimation), which chooses the lag order. Starting at lag 2 and proceeding to the maximum, it asks whether any coefficient in each lag block survives Benjamini-Hochberg at level alpha, stopping at the first block with none. Across a seven-point alpha grid from 0.01 to 0.30, it retains the choice with the lowest trailing validation RMSE. Finally, BH retains edges at q = 0.05. The theory makes the debiased coefficients asymptotically normal around a weighted average of drifting coefficients within each rolling window. It requires conditional homoskedasticity, a DML product-rate condition on the learners, and windows no longer than order T^(2/3).
The synthetic experiment uses a three-variable system, T = 3,000 and 2,595 rolling windows. With LightGBM, ORACLE-VARX estimates the true lag order at RMSE 0.962, versus 1.475 for VAR and 1.214 for PCMCI. Its edge false discovery rate (FDR) is 0.047 against a 0.05 target; PCMCI records 0.045 and VAR 0.129. The market experiment uses daily open-to-close returns from 2000 to 2020 on XLY, XLP, XLE, XLF, XLV, XLI, XLB, XLK and XLU. Testing runs from 2004 to 2020, with ten FRED macro series, a 504-day window and up to 10 lags. For the headline weighted book, each ETF receives its forecast divided by the sum of absolute forecasts. SPY offsets net exposure, and the book rebalances daily.
Does the morning lag choice earn its keep?
On the authors' figures, it does. Their VAR baseline selects lags by trailing validation RMSE and earns a 1.08 Sharpe (5.93% a year) on the weighted book. ACLE-VAR changes that lag rule to a significance gate and earns 1.32 (7.34%). Every reported strategy column improves: naive rises from 1.04 to 1.15, Top 50% from 0.95 to 1.31, Top 25% from 1.01 to 1.31, and Top 75%, trading only two ETFs, from 0.86 to 1.04. On synthetic data, ACLE cuts VAR lag RMSE by 32% (1.475 to 1.006).
The authors are explicit about the limit: "ACLE is a heuristic; we do not prove that it selects the true lag." Graphs lead their abstract, yet the contribution list assigns the roughly 1.3 Sharpe to "the ACLE adaptive lag procedure alone". The narrower trading claim comes from the paper itself.
DML costs the trade Sharpe
Adding DML weakens the ETF results. ORACLE-VARX with Extra Trees earns 0.83, 0.85 and 0.68 Sharpe for the vix, macro5 and all10 confounder sets. TabPFN earns 0.70, 0.74 and 0.36; LightGBM on all10 falls to 0.15. In the weighted book, linear inclusion of confounders lifts VIX alone slightly (1.11 vs 1.08), then cuts performance to 0.43 on macro5 and 0.75 on all10. The strongest DML result is 1.15, from Extra Trees on macro5 under Top 50%. It is the maximum in a 150-cell grid (5 strategies, 3 presets, 5 learners, 2 lag rules).
The authors give causal graph recovery priority over trading performance. In the main text, they report that ACLE-VAR produces the best portfolio without DML or confounders. They also acknowledge a serious violation of causal sufficiency in sector ETFs: many common factors are unobserved. In their account, hidden factors leave predictive links for a short-lag VAR to exploit, while DML adds parameters that must be fitted on 504 noisy days. An appendix says DML "trades predictive power for causal validity". The synthetic forecast results complicate that explanation. Most DML variants improve forecasting there, and TabPFN reaches MAE 0.091 against VAR's 0.110. The ETF shortfall has not been shown to be an inherent cost of debiasing.
The synthetic graph results deserve credit, with a narrow reading. LightGBM alone meets the 0.05 FDR target. TabPFN records 0.066; the other three learners range from 0.10 to 0.12. For Extra Trees (0.101), the authors suspect that first-stage accuracy falls short of their condition. The abstract's 0.047 therefore describes one learner of five. LightGBM's FDR moves from 0.047 with all three confounders to 0.084 with two and 0.180 with one, while PCMCI remains at 0.052. Testing all five candidate lags rather than only those selected brings FDR down to 0.023 and raises power slightly (0.676 vs 0.673). The adaptive cutoff costs some error control. A plot for ORACLE-VARX with Extra Trees and all ten confounders shows p̂ near 1 in 2004 to 2006, rising to 3 to 5 in 2008 and early 2020, against 21-day SPY realized volatility. We did not find a test statistic for that pattern.
The implementation gaps
The portfolio instructions are specific; the trading bill is omitted. "Transaction costs are not modeled." Ten instruments are rebalanced daily. SPY provides a dollar hedge, though the authors note that market exposure may remain. We did not find the five constituents of macro5 or a stated fill convention at the open.
ACLE uses classical OLS z-statistics for its stopping rule, without HAC correction. The paper's homoskedasticity assumption supports that choice. Daily sector returns put pressure on the assumption, especially when the paper's own lag estimates jump with volatility, and those z-statistics determine which lags survive. Cross-fitting folds have a zero buffer by design: the authors argue that drifting nuisance functions become stale.
Our Extra Trees run, net of commissions
We traded ORACLE-VARX using Extra Trees with VIX as the sole confounder. This DML configuration earns 0.83 Sharpe in the paper; its best trade is ACLE-VAR at 1.32. Our January 2020 to July 2024 run is net of $0.004 a share in commissions, subject to a $1 minimum per order. It returned -2.25% (CAGR -0.50%), with a -0.33 Sharpe and -4.59% maximum drawdown across 22,586 trades. Profit factor was 0.98; win rate was 48.42%. The SPY offset left beta at -0.01 in our run, so market drift does not account for the loss. For the same configuration, the authors report 0.83 Sharpe and 4.33% a year from 2004 to 2020. Their result is gross; ours is net. The windows overlap by about one year, and the figures measure different things.
Commissions plainly hurt our run. We averaged about 20 fills a day and capped each leg at 10% weight. That cap brought annualized volatility down to 1.50%, against roughly 5% implied by the paper's figures. At that size, 0.83 Sharpe amounts to only about 1.25% a year in gross edge. Per-order minimums across roughly 20 daily fills could eat much of it, though we cannot measure how much. A 0.98 profit factor is consistent with a near-zero gross edge falling below zero after costs.
The period likely matters too. Our 2021 to mid-2024 observations follow the paper's December 2020 data end. Sampling may explain part of the difference: with 1,131 daily observations, Sharpe standard error is about 0.47, leaving room for noise in the 1.16-point gap. We cannot account for the whole gap from what we can see. We also did not log selected lags, so the volatility pattern remains unchecked. This was one automated pass, evidence about our build first rather than a verdict on the authors' work.
An ACLE-VAR run net of costs on post-2020 data would change our view. ACLE-VAR produced the paper's 1.32, measured gross over 2004 to 2020.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.