A better VaR forecast is not yet a better trading decision. Ormaniec, Pitera, Safarveisi and Schmidt report a lower mean quantile score for their five-neuron LSTM than for all four benchmarks on 19 of 25 Fama-French portfolios in a test period ending March 2019. When COVID enters the test period, that count drops to 12, against 10 for t-GARCH. The paper measures forecasts alone. It sizes no positions, leaving the value of that score edge at a desk open.
The paper is "Estimating value-at-risk: LSTM vs. GARCH" (arXiv 2207.10539, 2022). GARCH puts volatility clustering into VaR; earlier neural approaches estimated GARCH or CAViaR parameters for use in a VaR formula. This LSTM learns the VaR quantile directly, with no assumed dynamics. The authors test it on 8 simulated GARCH processes and 25 Fama-French portfolios sorted on book-to-market and operating profitability, using two 7,500-day windows. They report no trading P&L and measure neither capital efficiency nor sizing efficiency.
What goes into the forecast?
For VaR 1%, the network takes a rolling window of 250 daily returns; for VaR 5%, it takes 50. After demeaning the window, it receives five channels per day. One carries the window mean. The others carry the first four Chebyshev polynomials of each demeaned return; 2x²−1 and 8x⁴−8x²+1 represent squared and quartic deviations. The authors use this orthogonal basis instead of raw moments to keep the second and fourth powers from cancelling each other during training.
One LSTM layer has 5 units. A 16-neuron ReLU layer follows, then a single VaR output. The training objective is the mean quantile score, (1{x≥y}−α)(x−y), which lets the network learn the conditional α-quantile without assuming a noise distribution. A GARCH recursion gives declining weights to past squared returns. Here, the LSTM sees those squares in sequence and can assign different weights, including weights that vary nonlinearly.
Each series supplies 7,500 daily observations in chronological 80/10/10 train, validation and test splits. The authors fit a model on the first 80%, tune it on the next 10% and score it on the last 10%. Of two calibrations, they retain the one with more consistent behaviour across training and validation.
The four benchmarks are the empirical quantile; a Gaussian estimator using a Student-t quantile and a √((n+1)/n) bias correction; a Gaussian GARCH with that same correction, introduced in this paper; and a t-GARCH. The GARCH models are refit for every window. For real data, the GARCH specification is GARCH(1,1) only.
The simulated advantage has a limit
At VaR 1%, the LSTM loses to GARCH in all eight specifications. At VaR 5%, it wins seven. These simulations cover eight GARCH(p,1) specifications, with p from 1 to 4 and Gaussian or t5 noise, run for 100 resamples each. At VaR 1% with n=250, GARCH(1,1)-n gives the LSTM a score of 2.79 against 2.74 for GARCH; true VaR scores 2.69. The GARCH(4,1)-n result is much worse: the LSTM scores 3.57 with a standard deviation of 3.77, against 2.57, and its exception rate reaches 1.75%.
At VaR 5% with n=50, the LSTM scores 10.56 against 10.73 in GARCH(1,1)-n, while true VaR scores 10.36, all in units of 10^-4. Its one loss is GARCH(3,1)-t, at 9.64 against 9.46. Across the 800 backtests at each level, it records the smallest score 32% of the time at 1% and 66% at 5%.
The authors call n=50 a small-sample test. GARCH gets 50 observations for each fit, while the LSTM has already trained on roughly 6,000 observations of the same process. That difference in available information bears on the 5% win as much as the choice of architecture does. There is an advantage running the other way, which the authors acknowledge: simulated GARCH receives the true p and noise family.
For the non-stressed real-data window, 02.06.1989 to 08.03.2019, the LSTM posts the lowest score in 19 of 25 portfolios. GARCH-t leads in 3, the empirical estimator in 2 and the unbiased estimator in 1. Exception rates tell a narrower story. Counting the rate nearest 1%, the LSTM and empirical quantile each lead in 14 portfolios; GARCH-t leads in 3.
Those exception-rate counts rest on few breaches. In a backtest of about 500 days, a 1% target implies 5 expected exceptions, and 1.40% means 7. One breach shifts the rate by 0.20 points over 500 days. We did not find Kupiec or Christoffersen coverage tests, or Diebold-Mariano score comparisons, in the paper. The reported evidence is bold counts and averages.
COVID changes the exception-rate leader
The second window covers 20.05.1992 to 28.02.2022, putting the March 2020 crash in its test slice. The LSTM retains a narrow score lead: 12 portfolios, versus 10 for GARCH-t and 3 for GARCH-n. On exception rate, the empirical quantile leads in 16, GARCH-t in 9 and the LSTM in 7.
HiBMHiOP is the LSTM's worst portfolio. Its score is 17.54 against 9.18 for GARCH-t, and it breaches on 3.40% of days against 1.40%. The excessive breach rate means it held too little capital on some days.
The authors describe the crisis response directly: "Not trained on such extreme scenarios, the LSTM estimator overshoots to a relatively high level in the pandemic crisis." GARCH "does not require those high levels of capital," they add, calling that advantageous. They argue that the LSTM "quickly reverts to GARCH-type dynamics." We accept the reversion shown in their plots, although a few paragraphs earlier the paper says the GARCH and LSTM estimators "cool down much slower". The overshoot remains consequential. We found no retraining described: one model per series is fitted on the first 80% and used throughout the test slice.
The conclusion claims that the LSTM "outperforms all existing estimators in terms of exception rate and mean quantile score" on market data. The tables give an exception-rate tie of 14-14 in one window and a loss of 16-7 to the empirical quantile in the other. Section 5.2 acknowledges both results. For the calm window, it says the LSTM "outperforms the other estimators in 14 cases, just as the empirical estimator does"; for the stressed window, "the empirical estimator clearly performs best (in 16 out of 25 cases), as we already expected". The conclusion conflicts with that results text.
The abstract at the head of the paper describes "calibration properties comparable to those of classical parametric benchmarks". The original arXiv text's abstract instead repeats the conclusion's outperformance line. The tables support the abstract at the head of the paper. They also support its claim that forecasts "frequently achieve lower average quantile scores, suggesting improved tail risk assessment": the counts are 19/25 and 12/25. The quantile-score edge holds in the calm window and is close in the stressed one.
Two items remain open. The paper supplies no VaR 5% table for either Fama-French window. It calls the results similar but shows VaR 5% paths only for two unnamed portfolios in the COVID window. Inputs are rescaled to [0,1], and we did not find which data were used to fit that scaling. Fitting it on the full series would leak the test-period range into training. The paper measures forecast loss only, so the capital or position-sizing value of a lower pinball score remains unmeasured.