Most of the apparent advantage from choosing a crypto volatility loss comes from shifting the forecast level. Tokajuk and Chudziak make that case convincingly. Their second claim, that model choice matters more after alignment, has weaker support. And the VaR table, used by the authors only to diagnose forecast level, leaves aligned Gaussian VaR breaching on 3.48% to 3.61% of days despite a 5% target.
Loss determines the level
Gneiting's elicitation argument supplies the mechanism. With squared log error, a model learns the conditional geometric mean of next-day variance. QLIKE elicits the conditional mean, while pinball on log error elicits the median. HMSE uses squared relative error and targets E[h²]/E[h]. That quantity always exceeds the mean, with the gap widening during spike regimes.
The paper applies this logic to crypto variance. Different losses can favour different levels when variance jumps, even when their forecasts follow the same daily path.
Seven losses are run across five models: HAR, HAR-X, NeuralHAR, DLinear and LightGBM. The sample contains daily Binance USDT spot bars for BTC, ETH, BNB, XRP and ADA through 31 May 2026. Next-day Garman-Klass variance is the target. Five expanding folds for each coin produce 25 blocks, each containing 253 validation days and 379 test days.
The comparison rests on a ratio. Its numerator takes the average absolute gap in log test QLIKE across the 21 loss pairs. The denominator applies the same calculation across the 10 model pairs. Before alignment, the median ratio reaches 2.91 (95% CI [2.27, 3.25]). Loss selection therefore appears almost three times as consequential as architecture plus inputs.
Alignment changes the picture. For each loss-model fit, the authors calculate one constant: validation mean Garman-Klass variance divided by forecast variance. They multiply every test forecast by this QLIKE-minimising rescale. Daily movements remain intact while the level shifts. Cross-loss variation falls 77% (CI [72%, 79%]), taking the ratio down to 0.67 (CI [0.56, 0.99]).
LevelShare offers another view. It measures how much of the squared log gap between forecasts trained with two losses can be explained by a constant offset. LevelShare reaches 89% for HAR and 56% for LightGBM.
The decomposition is clean, with an explicit target for every loss. Circularity is the obvious challenge because forecasts are aligned and scored with QLIKE. The paper addresses it directly. Under MSE-log and HMAE evaluation, the aligned ratio remains below one, and the same holds for every coin. Before alignment, those two scores produced ratios of 3.71 and 3.98.
The level argument convinces me.
How fair is the model contest?
The claimed reversal deserves more caution. Models retain their native inputs in the primary grid. DLinear receives 48 days of target history. HAR-X, NeuralHAR and LightGBM also receive returns, leverage, volume and alternative range proxies. For the altcoins, they additionally use lagged BTC variance. The authors acknowledge that "model choice" therefore includes some input choice.
Once inputs are shared, the aligned ratio becomes 0.81 with an interval of [0.66, 1.22]. The paper describes the reversal as statistically clear only for the full grid, while its conclusion calls the shared-input result inconclusive. The abstract likewise confines the claim to the full five-model comparison.
The disclosure is candid. Still, the shared-input grid provides the fair comparison, and it cannot establish whether architecture or loss matters more after level adjustment. A desk can also implement one validation rescale more cheaply than it can replace either the model or the loss. That rescale removed 77% of cross-loss QLIKE variation.
Interactions disappear inside the average ratio, another limitation the authors concede. HMSE-trained LightGBM records 0.659 QLIKE. The best cell, NeuralHAR with QLIKE, records 0.576, while LightGBM's own best loss, Patton-b, gives 0.625. HMSE-trained HAR reaches 1.243.
VaR converges and misses 5%
The raw 5% Gaussian VaR breach rates are:
- Pinball 6.31%
- Huber-log 6.23%
- MSE-log 5.92%
- QLIKE and Patton-b 3.68% each
- HMAE 2.19%
- HMSE 1.28%
Alignment compresses every loss into a range of 3.48% to 3.61%, eliminating 97% of the cross-loss spread. The authors interpret this as forecast level carrying through to risk estimates. The table supports that interpretation.
But the raw results deserve another look. Although the log-error losses perform poorly on QLIKE, their VaR estimates lie closest to nominal coverage. MSE-log, at 5.92%, is the table's best-calibrated VaR. Alignment moves it to 3.57%.
The adjustment shifts MSE-log from 0.92 points above nominal to 1.43 points below it. Too many breaches become too few, and all seven losses end up similarly over-conservative. Kupiec rejects correct coverage in 128 of 175 series before alignment and still rejects it in 109 of 175 afterward. The authors state the problem plainly: "convergence across losses does not imply correct VaR calibration."
The 3.5% breach rate
The authors propose three candidates: the Gaussian assumption, the variance proxy and forecast error. I place more weight on the Gaussian assumption, together with the constant's inability to follow time-varying bias, which the authors concede.
If variance is correctly estimated while return tails are fat, a Gaussian 5% quantile overshoots. A heavy-tailed density matched on variance places more mass near the centre, so the 5% point for a unit-variance Student-t with few degrees of freedom falls inside 1.645. The rescaling constant creates a separate effect. Because it averages h/f, a handful of spike days within the 253 validation days can raise the multiplier across all 379 test days. Both mechanisms push breach rates below 5%, matching the paper's direction.
The authors also observe that alignment optimizes QLIKE rather than tail coverage. QLIKE calibrates mean variance. VaR depends on the 5% return quantile.
Our extension would add a single rescale chosen to hit 5% breaches during validation, then apply it in the same fashion. This would help separate level effects from distributional shape. The paper already lists non-Gaussian risk models as future work. If a quantile rescale still left Kupiec rejecting across most of the 175 series, I would abandon the Gaussian explanation.
We could not test these explanations ourselves. Our crypto price history spans roughly 2018 to 2024 and therefore stops short of the paper's sample through May 2026. Our bars may differ from Binance USDT spot, while BNB coverage remains unconfirmed. Every experimental figure reported above belongs to the authors.
The paper is worth reading for its decomposition. Raw and aligned results should be reported side by side, as the authors recommend, while the VaR table should be treated as the start of a calibration problem.