A desk should test George Panos's decision matrix before using its Transformer cutoff to choose a model. The review puts the minimum at 10,000 to 50,000 observations, above two Transformer wins it cites. Its real-time recommendation also changes meaning as the discussion moves from inference to retraining.
The case for each model
Panos asks why LSTMs win some volatility forecasts and Transformers win others. An LSTM carries a gated cell state forward, one step at a time, using the same weights across time. The paper credits that design with filtering noise and working on 500 to 1,000 observations. It puts effective memory at 200 to 500 steps. A Transformer allows each time step to attend directly to every other. The paper credits it with longer memory and faster resets after regime shifts, while assigning it a cost of O(T²·d) and a need for 10,000 to 50,000 observations.
The paper states: "We do not conduct new empirical backtesting or original coding." It derives five predictions from the architectures and compares them with RMSEs quoted from six studies. Every figure below is an RMSE from those studies. None is Panos's own result, and the review supplies no model or code.
On daily VIX, LSTM wins 2.10 to 2.28 with 3,528 observations (p=0.12, not significant at 95%). With 5,040 observations it wins 2.05 to 2.18; a hybrid, which puts an attention layer on LSTM hidden states, reaches 1.96, significant at p<0.05. At 7,812 observations, Transformer moves ahead, 1.98 to 2.01.
Three VIX samples, one flip.
For hourly Bitcoin, Transformer wins 0.78% to 0.91% on about 26,000 observations. In a sample-size sweep, LSTM leads below 2,000 (1.02% vs 1.15%). The models tie at 2,000 to 5,000 (0.95% vs 0.94%), and Transformer leads above 8,000 (0.76% vs 0.89%). On clean Ethereum data, LSTM loses 0.80% to 0.72%; after simulated bid-ask bounce, it wins 0.95% to 1.15%. Panos's matrix accordingly favors LSTM for small, noisy or real-time problems, Transformer for long hourly histories in liquid venues, and the hybrid between them.
Which Transformer cutoff?
The review offers six thresholds. Its abstract and comparison table give 10,000 to 50,000. Section 3.3 attributes a 10,000 floor to Patel and Shah, although their reported win begins above 8,000. Prediction 3 has LSTM winning below 5,000 and Transformer catching up above 20,000. The VIX synthesis places the advantage above 5,000, the crypto synthesis above 8,000, and the closing checklist asks for more than 20,000.
The cited results strain the headline 10,000 minimum. Li et al. report a Transformer win at 7,812 observations, below the floor. Patel and Shah put their crypto win in a bucket beginning above 8,000, also below the stated floor at its lower end. Chen and Ge's VIX sample of 5,040 clears the review's 5,000 line, yet Transformer loses by 0.13. Li's winning margin is 0.03 RMSE, with no significance test reported. Patel and Shah label their T > 8,000 bucket ">12 months". The review also gives 8,760 hours per year, which makes 8,000 hours about eleven months.
The limitations section names five weaknesses in the source studies, including publication bias and data snooping, then argues that theory and evidence converge anyway. Prediction 1 is a problem for that claim. It expects Transformer to beat LSTM on VIX because the Hurst exponent is near 0.65. Section 4.1.2 adds a sample-size condition: once T exceeds 2,500, Transformer should match or slightly exceed LSTM. It loses at 3,528 (2.28 vs 2.10) and at 5,040 (2.18 vs 2.05). Both samples pass that threshold. The VIX synthesis nevertheless says the evidence supports the predictions, locates a Transformer edge only above 5,000 observations, and favors LSTM in practice because few desks have 20 years of data.
Some cited sources warrant checking before anyone sizes a position from them. Patel and Shah appears as arXiv:2303.12345, an identifier that looks like a placeholder; verify it before relying on its figures. Liu et al. (2019), cited as a Journal of Financial Data Science LSTM and Transformer comparison on VIX, calls for the same check.
How many regimes fit in an hourly sample?
An hourly bar adds less independent training information than the observation count suggests. The paper estimates 10 to 20 crypto regime shifts a year. At those rates, an 8,000-hour history, about eleven months, contains roughly 9 to 18 transitions. Li et al.'s 31 years of daily VIX contain roughly 15 to 30, using the paper's estimate of 5 to 10 shifts per decade.
Those counts put the supposedly large crypto sample near the sufficient VIX sample in exposure to regime changes. For crypto, the paper attributes Transformer's edge partly to attention resetting after shifts, alongside sample size. If resets drive the result, shift count deserves more weight than raw hours; an 8,000-hour sample contains roughly 9 to 18 shifts. Wang et al. use a 1,000-hour sliding window, which spans one or two shifts at a time on the paper's rates. T alone cannot tell us whether that model learned to generalise across regimes or learned the latest one.
Real-time inference and retraining
Prediction 5 compares O(h²) per LSTM step with O(T²) for Transformer and frames the step as inference. Section 3.3 describes online operation as updating after each new observation "without full retraining". Its O(h²) step updates an LSTM hidden state while weights stay frozen. Refitting the weights costs O(epochs × T × h²) under the paper's training formula. Yet the VIX section calls its online case "daily retraining" and still marks LSTM strongly preferred.
The comparison needs separate clocks. Freeze both models' weights and time each forecast, using a sliding window for Transformer. Then refit both on a schedule and measure the hours each needs to recover forecast error after a shift. The matrix calls online Transformer hard, while Section 5.4 puts a 1,000-step sliding window at about 1ms and reserves LSTM for microsecond budgets. The microseconds-versus-seconds distinction gives a desk a usable boundary: take LSTM under a sub-millisecond budget; with seconds available, the paper's timing leaves a windowed Transformer in contention.
A Patel-style sweep reporting crossover by regime transitions rather than hours, with a significance test at each step, would change my mind.