AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

ALM-GARCH Rejects Symmetry Six Times and Still Loses to FIAPARCH

Positive FTSE returns add 4.6e-7 of variance mass; memory survives in three of six series

2026-09-09 · 9 min read · US equities/ETFs and crypto

Reviewing: Asymmetric Long-Memory GARCH: Sign-Dependent Kernel Injection in a Two-Dimensional Markov Chain · Kennedy Titus Kayaki and Kyungsub Lee · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Sign-Dependent Long-Memory ALM-GARCH Volatility Targeting

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

Run on a different market than the paper

The paper estimates volatility dynamics on several international equity indices plus Bitcoin. We would apply the same return-volatility forecasting mechanism to US equity ETFs or liquid US equities, while retaining Bitcoin/crypto where available; sign-dependent volatility responses and long-memory variance dynamics are price-return mechanisms that can survive this market substitution. Reported forecast results for the original international-index universe do not transfer to the US universe.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (10)

  • deviation left undescribed by the audit (invalidates: The paper's reported Table 7 QLIKE levels and Diebold-Mariano significance results; direct replication of the paper's expanding-window forecast experiment)
  • deviation left undescribed by the audit (invalidates: The paper's final-30-percent out-of-sample sample composition; direct comparison to the paper's market-specific average QLIKE values)
  • deviation left undescribed by the audit (invalidates: The paper's Table 7 QLIKE values; the reported HAR-RV ordering; the reported market-specific Diebold-Mariano test outcomes; direct replication of the paper's five-minute and one-minute realized-variance evaluation)
  • deviation left undescribed by the audit (invalidates: All paper cross-market rejection classifications; all paper asset-level persistence estimates; all paper asset-level QLIKE and likelihood results)

6 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 7.4%Jul 2024
Sharpe
0.19
Total Return
7.4%
Max Drawdown
-20.9%
CAGR
1.6%
Volatility
8.6%
Beta vs SPY
0.29
Trades
5,575

The FTSE 100 fit assigns almost the entire variance response to negative returns. Its positive-shock injection amplitude is 4.592e-7, beside 0.202 for the negative branch. The S&P 500 looks much the same at 1.654e-5 versus 0.201. Taken literally, positive returns add almost no variance mass on either index. These estimates turn the leverage term from a small adjustment into nearly the whole response.

The rest of the paper asks whether that literal reading can be trusted.

The model Kayaki and Lee built

Lee and Kayaki begin with their earlier symmetric long-memory model. It replaces an infinite-order fractional filter with a two-dimensional Markov chain. The X coordinate tracks accumulated variance. A second coordinate, c, records an effective kernel age, compressing the arrival times of past shocks into one state variable.

Each shock contributes a power-law kernel with amplitude xi and offset gamma. Between arrivals, the existing kernel ages deterministically. This construction generates hyperbolic decay without retaining the full return history. Conditional variance is sigma-squared = mu + X and remains predictable one step ahead.

ALM-GARCH allows shock sign to determine the amplitude and offset. The amplitude difference, xi_minus versus xi_plus, forms the level channel and recasts the familiar leverage effect. The offset difference, gamma_minus versus gamma_plus, forms the memory channel. Sign changes the age at which the kernel resets, altering its finite-horizon persistence profile. Both branches retain a common power exponent p and therefore the same asymptotic decay order. The authors define memory asymmetry through kernel age, rather than through separate tail exponents. For the direct fit comparison, ALM-GARCH estimates five parameters with p = 1.2 fixed; FIAPARCH estimates six.

The equity data cover five Oxford-Man indices. Returns are daily open-to-close, realized variance comes from 5-minute observations, and the period runs from January 2000 to 27 June 2018. The samples are S&P 500 N=4641, FTSE 100 N=4660, DAX N=4692, Nikkei 225 N=4501 and KOSPI N=4550. Bitcoin uses Bitstamp UTC close-to-close returns and 1-minute realized variance from January 2012 to January 2025, N=4754.

The conditional mean is fixed at zero. Estimation uses Gaussian quasi-ML, with p fixed at 1.2 and xi fitted on the log scale. Likelihood-ratio tests for the channels use a 999-replication bootstrap, except for the DAX memory cell, which uses 3,999.

Positive Harris recurrence is also proved for interior configurations. The argument applies a Foster-Lyapunov drift with V(x,c) = log(1+x) + lambda*c. Geometric ergodicity and score asymptotics are absent. The authors consequently use chi-square references as calibrations rather than inference.

How much survives the power curve?

Every series rejects joint symmetry. LR_sym is 187.09 for the S&P 500, 210.87 for the FTSE 100, 157.66 for DAX, 38.38 for Nikkei, 49.86 for KOSPI and 25.83 for Bitcoin. Every bootstrap p = 0.001. The level channel supplies most of that result: all six reject, with LR_lev ranging from 16.62 for Bitcoin to 110.01 for the S&P 500. A symmetrized-residual bootstrap removes the Gaussian innovation assumption and produces p values between 0.001 and 0.035.

Memory asymmetry has a thinner record, as the authors acknowledge. The bootstrap rejects for Nikkei 225 at LR 15.01, KOSPI at 26.08 and Bitcoin at 24.82. Under the symmetrized-residual procedure, Nikkei's memory p is 0.046, marginal at 5%. KOSPI comes in at 0.001 and Bitcoin at 0.011.

DAX misses, with 7.23 and bootstrap p = 0.066 from 3,999 replications. The S&P 500 also fails to reject at 0.15, p = 0.853, as does the FTSE 100 at 0.95, p = 0.573. Those last two results need care. Their fitted positive amplitudes, 1.654e-5 and 4.592e-7, fall below the paper's pre-specified 1e-4 tolerance. An offset cannot be identified on a positive branch lying that close to zero, so the non-rejections do not establish equal offsets.

Where identification survives, the usual ordering reverses. Positive shocks contribute less variance and then decay more slowly. Nikkei's fitted one-step persistence ratios are rho_plus = 0.999 and rho_minus = 0.733. Bitcoin gives 0.875 against 0.528.

The paper's Monte Carlo shows why the channel labels deserve restraint. Near the null, local scores for the two channels have a correlation of about 0.95. At N = 4,000, power reaches 0.302 for the level test, 0.199 for memory and 0.226 for the joint test, each at nominal 5%. The authors therefore place more weight on the joint test and qualify the separate channel results. Their own figures support that choice.

Conventional chi-square-1 calibration fares badly at several fitted memory-restricted nulls. It rejects a true null 24.6% of the time for the S&P 500, 19.7% for the FTSE and 22.0% for DAX. At the Nikkei, KOSPI and Bitcoin nulls, rejection rates are 5.7%, 5.6% and 5.8%. Without calibration, DAX changes from non-rejection to rejection: LR 7.23, chi-square p = 0.007, bootstrap p = 0.066.

Anyone applying nested LR tests to a near-boundary variance parameter should spend time with that table.

FIAPARCH takes every fit comparison

FIAPARCH(1,d,1) beats ALM-GARCH in all six series by quasi-log-likelihood and BIC. The likelihood advantages are +56.2 for the S&P 500, +17.0 for FTSE, +20.9 for DAX, +38.5 for Nikkei, +37.6 for KOSPI and +24.6 for Bitcoin. BIC differences range from -25.5 to -104.0 in FIAPARCH's favour, despite its six estimated parameters against ALM-GARCH's five.

The paper says this plainly in the Introduction, which "favors FIAPARCH by BIC in all six series", and repeats it in the Conclusion: "so ALM-GARCH does not provide a general fit advantage". Its answer follows at once. ALM-GARCH "instead provides a separately testable distinction between the level and memory channels." The abstract uses softer language, "Out-of-sample performance is broadly comparable to standard benchmarks", while leaving FIAPARCH unmentioned.

Out-of-sample evaluation uses the final 30% of each series. Forecasts are one-step ahead from an expanding window, with refits every 250 days. QLIKE is measured against realized variance. Diebold-Mariano tests use Newey-West lag floor(n^(1/3)).

HAR-RV records the lowest QLIKE in every market. For the S&P 500, it scores 0.245 against 0.363 for ALM-GARCH. KOSPI is 0.141 against 0.207. DM statistics versus HAR range from +2.58 to +5.79 and are significant at 1% in five of six cases. FTSE is the exception at t = 0.62. HAR forecasts realized variance from its own lags, giving it a different information set, a distinction the paper makes explicitly.

Results against return-based models are scattered. ALM-GARCH significantly beats GARCH(1,1) only for DAX, DM -2.11. It beats GJR for DAX at -3.60 and KOSPI at -4.98. EGARCH wins for the S&P 500 at +3.05 and Nikkei at +3.51, while FIGARCH wins for KOSPI at +6.15 and Bitcoin at +4.81. Against its symmetric LM-GARCH predecessor, ALM-GARCH achieves a significant improvement in one market of six. DAX has QLIKE 0.203 versus 0.221, with DM -2.03.

The exponent remains another weak point, and the paper flags it. Across p in [1.05, 3.0], maximized log-likelihood changes by no more than 8.4 points. On the equity indices, the persistence summary rho shifts only 0.05-0.08. When freely estimated, p reaches the imposed boundary for several assets. Fixing p at 1.2 costs at most 6.6 points. The authors treat p as an identifying normalization rather than a fitted shape parameter, adopting 1.2 from the symmetric base model to preserve comparability. Our observation is that this choice and the in-sample rankings use the same full sample as the channel tests.

Stability also comes close to the edge. The fitted Nikkei configuration fails the certificate with M = +0.003054, while KOSPI is borderline at M = -0.000088. The authors refit under M* <= -1e-3. This restriction costs 0.61 log-likelihood points for Nikkei and 0.05 for KOSPI. Every channel LR moves by under 1.3 points, with no classification changed. The unconstrained solution lies almost on the frontier.

Our ETF adaptation

The paper evaluates QLIKE against realized variance. It runs no strategy or portfolio and includes no cost accounting. Its reported results therefore contain no strategy Sharpe or strategy return to set alongside the figures from our run. Every number below is ours alone.

We could not trade the paper's universe. We substituted five equity ETFs, SPY, EWU, EWG, EWJ and EWY, for the five international indices and omitted Bitcoin. This is an adaptation using US-listed ETF proxies, rather than a replication. We judged that sign-dependent variance response and slow variance decay could carry through the ETF wrapper. That judgment is ours and is not a paper finding. The paper's index QLIKE results do not transfer to these ETFs, and we make no comparison between them.

Our data are daily adjusted close-to-close returns from 2020-01-01 to 2024-07-01. Gaussian QML runs on a rolling window of up to 1,000 returns, with a minimum 750, and refits every 250 trading days. The two-dimensional state updates daily between refits. We set p = 1.2, tau = 1 and the conditional mean to zero.

One-day variance forecasts drive covariance normalization for equal base weights, targeting 10% annual volatility. Signals are formed after the close and executed at the next observed open. The portfolio rebalances daily with a zero no-trade threshold. We charge commissions of $0.0040 per share and a $1.00 minimum for every filled order. Short borrow, margin financing and market impact remain unmodelled. There is no leverage cap.

The portfolio returned 7.35% over the window. Sharpe was 0.19, Sortino 0.26, Calmar 0.08, max drawdown -20.94% and annualised volatility 8.59%. A Sharpe of 0.19 at 8.59% realized volatility is weak. With no gross cap, a 10% target and five equal-weight ETFs, the normalizer increases gross exposure during calm periods. Financing for that borrowing is unmodelled.

This run is one automated pass constructed from the paper's description. It does not isolate the sign-dependent kernel. A volatility target responds to the forecast level and is almost indifferent to the branch that generated it. The run contains no test of the memory channel, although that channel constitutes half of the paper's claimed contribution. Its weak performance speaks first to our overlay design and does not bear on the authors' statistical claims.

What remains useful

The decomposition and boundary diagnostics carry the paper. Kayaki and Lee show that making the injected kernel's offset sign-dependent allows positive and negative shocks to vary in persistence as well as size. The data support this memory distinction in three of six series. They also document over-rejection by the same nested test at the three boundary-prone nulls, S&P 500, FTSE 100 and DAX. Rejection rates there run from 19.7% to 24.6% against nominal 5%.

We have previously objected to volatility model rankings that omit a stated target or a Diebold-Mariano test (our note on nine FX series and seven models). This paper specifies its target and applies DM tests against six benchmarks. It also publishes the adverse comparisons: HAR-RV delivers the lowest QLIKE in all six markets, while FIAPARCH delivers the lower BIC in all six.

A memory-channel rejection in a market with an active positive branch, accompanied by an out-of-sample improvement over LM-GARCH, would change our reading. DAX currently supplies the sole significant forecasting gain over the symmetric model. It is also the sole series with an identified positive branch where the memory channel misses rejection at 5%, with bootstrap p = 0.066. The empirical case rests on those results pulling in opposite directions.

Read the paper for sign-dependent kernel age and its boundary size table. Keep FIAPARCH in place.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each ETF:
  1. Build adjusted close-to-close log returns using only observations dated on or before t.
  2. Once at least 750 valid observations exist, estimate the ALM-GARCH parameters
     by Gaussian QML on a rolling window of at most 1,000 returns.
  3. Re-estimate parameters every 250 trading days; update the two-dimensional
     state every day between fits.

For return η_t:
  predictable variance before observing η_t: h_t = μ + X_{t-1}
  shock magnitude: z_t = η_t²
  branch s = plus if η_t ≥ 0, otherwise minus
  r(c_{t-1}) = (1 + 1/c_{t-1})⁻¹
  X_t = X_{t-1} · r(c_{t-1})^p + ξ_s · z_t
  X_t/c_t = (X_{t-1}/c_{t-1}) · r(c_{t-1})^(p+1)
              + (ξ_s/γ_s) · z_t
  preserve both state coordinates and forecast h_{t+1} = μ + X_t

After date-t close:
  4. Retain instruments with finite model and covariance inputs.
  5. Assign each eligible instrument an equal base weight of 1/N_t.
  6. Covariance-normalize the eligible base vector to target 10% annual
     portfolio volatility; apply no position cap, smoothing, or separate
     gross-leverage cap.
  7. If any target differs from its pre-trade weight, submit the target for the
     next observed trading-day open. Skip an order if that open is missing,
     non-finite, or non-positive.
  8. Apply new weights only to returns occurring after execution and include
     the cash-weight change when measuring turnover.

Separately evaluate h_{t+1} against the next adjusted close-to-close squared
log return on common dates where every compared forecast and target is finite
and strictly positive. Run p = 2.0 only as an independent sensitivity, without
sharing states or parameters with the p = 1.2 baseline.