AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

DSTNet clears persistence on MAE; equity timing remains unproven

MAE improves by 0.7 to 4.5 percent; our SPY, QQQ, IWM, DIA, GLD book earned Sharpe 0.46

2026-10-08 · 8 min read · ML-based multi-horizon forecasting and tactical allocation · US equity-index ETFs and a gold ETF

Reviewing: DSTNet: Dynamic Spectral Trajectory Network for Causal Multi-Horizon Financial Forecasting · Aashish Bohra and Lokendra Vishwakarm · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

DSTNet Spectral-Trajectory Forecasts and Next-Session ETF Allocation

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

Run on a different market than the paper

Trade US-listed index-tracking ETFs such as SPY, QQQ, IWM, and DIA instead of index levels, and GLD instead of gold itself. The method forecasts patterns in daily price paths, which can be tested on these instruments, but this does not replicate results for the paper's full index universe.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (2)

  • Algorithm 7, PAPER allocation diagnostic: b=mean predicted return on the validation set; z_t=ŷ_1^(t)−b; if |z_t|<η*, π_t=π_{t−1}; else if z_t<−θ_off*, π_t=max(e_0*−λ*·|z_t|/θ*,e_floor*); else if θ_on*>0 and z_t>θ_on*, π_t=min(e_0*+λ*·|z_t|/θ*,e_max); otherwise π_t=e_0*; π_t=clip(π_t,e_min,e_max); net P&L_t=π_t(p_{t+1}−p_t)/p_t−κ|π_t−π_{t−1}|. These are the paper’s same-close diagnostic equations, NOT the task’s next-session ETF execution rule.: Retain the equation as a non-executed reference; trade a validation-thresholded, long-or-cash ETF signal only at the next available adjusted open, accounting for entry and exit costs. (invalidates: Paper equity allocation Sharpe and static 65% exposure findings as ETF performance predictions; paper NYSE allocation Sharpe; paper gold-proxy allocation Sharpe and cumulative-return comparisons)
  • Compute indicators AFTER causal OHLCV denoising; use training-only thresholds and indicator Z-scores at every expanding fit. Do not fit transforms on full test history or use bilateral, centered wavelets.: Preserve the order, causality, and training-only fits, but use the guaranteed dollar-volume field as the fifth channel instead of assuming share volume; choose stated trailing periods for two indicators whose periods were not supplied. (invalidates: Numerical replication of the paper's 32/32 seed-42 panel; paper persistence-improvement percentages; paper ablation percentages; paper seed-stability counts)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 21.1%Jul 2024
Sharpe
0.46
Total Return
21.1%
Max Drawdown
-24.7%
CAGR
4.4%
Volatility
10.4%
Beta vs SPY
0.34
Trades
366

What the paper reports for its own strategy

  • S&P 500 DSTNet asymmetric risk-off allocation, test 2023-06-01 to 2024-05-31, 10 bps per unit turnover, idealised close-to-close execution: Sharpe 2.15 (annualised by sqrt 252), MaxDD 6.76%, CumRet 16.44%, exposure 65.00% (buy-and-hold Sharpe 2.15).
  • NYSE DSTNet risk-off, same period, 10 bps: Sharpe 1.80, MaxDD 5.65%, CumRet 11.31%, exposure 56.77%, total turnover 12.21.
  • KOSPI DSTNet risk-off, same period, 10 bps: Sharpe 0.48, MaxDD 9.68%, CumRet 4.39%.
  • NASDAQ DSTNet risk-off, same period, 10 bps: Sharpe 1.72, MaxDD 8.08%, CumRet 18.02%.
  • Gold (GC=F) DSTNet long-only, same period, 10 bps: Sharpe 1.99, MaxDD 3.34%, CumRet 12.51%, exposure 47.83%, total turnover 6.11.
  • Gold DSTNet long-short, same period, 10 bps: Sharpe 1.80, MaxDD 0.58%, CumRet 2.11%, exposure 7.65%.

A forecast that beats yesterday's close can still be a poor equity trading signal. Bohra and Vishwakarma show both sides of that problem with DSTNet. Nine tuned forecasters, including PatchTST, iTransformer and N-BEATS, have mean absolute error 2.88% to 12.83% above yesterday's close. None wins a one-, three- or five-day cell. DSTNet does win, though its one-day edge is less than one percent. The authors are candid about where the evidence ends: at forecast error.

A causal filter bank, then two attention streams

The paper takes aim at how wavelet hybrids handle financial series. A two-sided convolution reads past the forecast origin. Other designs use the transform for denoising alone, or give the model only a spectral snapshot from the forecast origin. DSTNet instead tracks how that spectrum moves through time.

Each OHLCV channel first passes through a causal Symlet-4 denoiser, with thresholds fitted only on training folds. Seven indicators follow: RSI-10, Stochastic %K, CCI-14, OBV, ATR-14, Williams %R and ROC-12. They are Z-scored using training statistics. A one-sided filter bank derived from the Morlet wavelet then processes them at eight scales, {1,2,3,5,8,13,21,32}. Magnitudes across a 20-day lookback make a 20x8x7 tensor. The authors call it the Dynamic Spectral Trajectory: a record of the spectrum's path across all 20 days. Because the longest filter has 192 taps, they discard the first 192 samples of every series.

One attention stream reads across the 20 days; another reads across the 8 scales. Together they form what the authors call a factorized Scale-Temporal Spectral Transformer. A CNN-BiLSTM reads the denoised prices, and a learned two-token softmax gate mixes its output with the spectral branch. Four per-horizon gates produce cumulative log returns for 1, 3, 5 and 10 days in one pass. The data cover seven equity indices plus gold futures from Yahoo Finance, 2010 to May 2024. An untouched 2023-06-01 to 2024-05-31 hold-out supplies 235 to 246 test windows per series. All nine learned baselines were tuned with the same five-fold expanding walk-forward.

The paper claims no source of alpha. It calls the trading test a falsification diagnostic.

Can it beat yesterday's close?

Persistence, the zero-return forecast, is the serious competitor. On MAE and MAPE, it has the lowest non-DSTNet point estimate in 29 of 32 series-horizon cells. Across the panel, learned-baseline MAE ranges from +2.88% for vanilla Transformer to +12.83% for XGBoost relative to persistence. DSTNet comes in 2.27% below it. Every series shows a wider margin at longer horizons: 0.67-0.92% at one day, rising to 3.43-4.49% at ten. The authors confine this finding to absolute-error measures. Under RMSE, learned baselines beat persistence in 16 cells, including 13 at ten days.

We have seen deep models struggle to clear persistence before (our note on a bank contagion graph that only ties persistence). DSTNet clears it in every primary-seed MAE cell.

The ablations give the trajectory much of the credit. Replacing it with a static final-day slice increases panel-mean MAE by 5.79%, the largest deterioration. Equal weights in place of the learned gate cost 2.94%; removing the burn-in costs 0.92%. Joint attention across all 160 time-scale tokens moves the MAE point estimate by just 0.05%, suggesting the factorized version saves compute without a visible accuracy penalty. That joint variant was width-comparable rather than parameter-matched, and trained at batch size 16. The paper also notes that the No-DST intervals overlap those of the full model.

The margin under testing

At one day, the paired test against persistence cannot settle the ranking. The abstract says testing "is inconclusive against persistence and the strongest learned forecasters." Diebold-Mariano statistics against persistence range from 0.16 to 0.24. Paired tests in that same one-day account favour DSTNet over weaker learned baselines, reaching 2.43 to 4.43 against LSTM. Those baselines are furthest above persistence. Against DLinear, N-BEATS, PatchTST and iTransformer, just 4 of 32 comparisons reach an unadjusted 10% level.

The authors' preceding abstract sentence rests its case on consistency: DSTNet is "the only model below it in every cell." Their seed rerun makes that claim less secure. Across seeds 42, 101 and 202, seed-mean wins over the best fixed competitor drop to 26/32 for MAE and MAPE, and 25/32 for RMSE. Wins on all three seeds total 23/32 for MAE and RMSE, and 24/32 for MAPE. The worst cellwise MAE coefficient of variation is 10.60%; the ten-day margin reaches at most 4.49%. Only DSTNet was rerun, as the paper discloses.

There is no corresponding test for the larger long-horizon gains. Inference ends at one day, and a year contains about two dozen disjoint ten-day blocks. For S&P 500 at ten days, the per-horizon test MAE table reports 91.94 for DSTNet and 95.36 for persistence. Their 95% moving-block bootstrap intervals are [76.91, 110.95] and [79.58, 114.35].

The widening margin leaves me wary of drift. A small positive drift term could gain ground against a zero-return forecast as the horizon lengthens. In the paper's allocation table, S&P 500 buy-and-hold returned 26.09% during the hold-out year. The paper does not test that reading, and we did not find a historical-mean forecast among its ten competitors. I would want that benchmark over more than one year before putting much weight on the ten-day gain.

A causal filter, an idealised fill

The filter's causality claim holds. It sums over non-negative lags, discards the 192-sample zero-padding transient and refits denoising thresholds within each training fold. According to the paper, the wavelet-finance studies it reviews leave causality at the forecast origin unaddressed.

The allocation test uses a more generous trading convention, disclosed by the authors. A signal containing the close at day t takes a position at that same close, omitting the time needed to observe the bar and compute the forecast. Gold trades through GC=F; rolls and margin are ignored, also openly.

Six equity books barely move

For KOSPI, S&P 500, Nikkei 225, DJI, DAX and NASDAQ, the risk-off rule trades at the open of the test and then stays put. Turnover totals 0.65, leaving a static 65% position. Its Sharpe consequently matches buy-and-hold by construction: 2.15 for each on S&P 500. The authors call this a scaling identity.

The regime diagnostics show how little direction the signal supplies. After subtraction of the validation-mean forecast, the S&P 500 one-day test signal gets 98.13% of down days right and 2.22% of up days right. Youden's J, sensitivity plus specificity minus 100, comes to +0.35 percentage points. In a year when the market rose 26%, a forecast negative nearly every day offered almost no discrimination.

NYSE is the equity exception that trades. Its average exposure is 56.77%, and drawdown declines from 10.66% to 5.65%. Sharpe declines too, from 2.02 for buy-and-hold to 1.80 for the risk-off rule. Cumulative return drops from 22.44% to 11.31%.

Gold gives the rule its best case, with the panel's highest J at +8.40. After 10 bps per unit of turnover, the long-only rule earns Sharpe 1.99, versus 1.74 for buy-and-hold and 1.71 for the volatility-targeted version. Drawdown falls from 8.16% to 3.34%. At 47.83% average exposure, cumulative return is 12.51% against buy-and-hold's 21.88%. The long-short variant brings drawdown down to 0.58%, with 2.11% return on 7.65% exposure. This is one series over one year, selected from a validation-searched grid, without paired inference. The paper treats it as a reason to test further. So do we.

Our ETF run: Sharpe 0.46

We could not trade index levels or gold futures, so we substituted SPY, QQQ, IWM, DIA and GLD, training one DSTNet per ETF. This adaptation came from one automated pass based on the paper's description. Its rule differs as well: each ETF takes a 20% long position when its bias-corrected one-day forecast exceeds a validation-selected threshold, and otherwise holds cash. All figures in this section are ours. They cover 2020-01-17 to 2024-07-01 and are net of per-share commissions.

Our book returned 21.14% across 366 trades, with Sharpe 0.46 and maximum drawdown -24.69%. The paper reports 2.15 for its S&P 500 risk-off rule and 1.99 for its gold long-only rule. Both exceed our 0.46, though the S&P 500 figure is buy-and-hold beta in a bull year. Our beta to SPY was 0.34; we could not have captured that rally in full. The paper's figures come from a different rule, a one-year window, and index and futures markets. Their distance from ours tells us little about the paper's result.

The paper's tables help explain part of the gap. In its strong uptrend year, S&P 500 buy-and-hold itself delivered Sharpe 2.15. Our period contains the 2020 crash and the 2022 bear market, where a book capped at 100% gross suffered that -24.69% drawdown. The paper's equity rule also kept static exposure on six of seven indices. Ours switched between 20% long and cash for 366 trades, using a signal whose equity J the paper finds near zero.

Other differences may matter. IWM is a small-cap fund without a counterpart among the paper's seven indices. We did not use the paper's same-close fill, and cannot isolate the cost of that change. Minimum-ticket commissions on 20% positions may have added drag. Our run speaks to our setup alone; one automated pass is no verdict on the authors' work.

What remains worth testing

The causal trajectory is the contribution I would keep. Removing it raises panel-mean MAE by 5.79% in the authors' ablation, though that is a point estimate and the No-DST intervals overlap the full model's. Their other useful result is the table in which nine learned forecasters, from LSTM to PatchTST, lose to a zero-return forecast on mean absolute error at one to five days. Their allocation test finds no equity timing value, and our one-pass ETF adaptation found none either.

The authors propose comparing persistence at three, five and ten days using stored errors, without retraining. Even a significant ten-day result against persistence would leave the drift question open. I would wait for a ten-day test against both persistence and a historical-mean forecast, across more than one year.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each ETF, build trailing features from adjusted daily OHLC and dollar volume.
Apply causal denoising and construct a 20-session, eight-scale spectral trajectory.
At each fit, train a separate DSTNet on eligible chronological samples whose four return targets are known.
On the preceding completed validation segment, estimate the one-day forecast bias and select a net-Sharpe-maximizing threshold; otherwise hold cash.
After session t closes, subtract the bias from the h=1 forecast. Target 20% long if it strictly exceeds the threshold; otherwise target cash.
Under the specified rule, execute changes at the next session's adjusted open and charge turnover costs, including initial entries and exits.

The supplied code is truncated, so the full training, calibration and order path cannot be verified from it.