Across these 16 metals, the wrong model can be costly: vanadium's QLIKE loss rose 345.99% over GARCH(1,1) with the wrong stochastic-volatility variant. Bastianin, Li and Shamsudin show how differently volatility behaves across sixteen energy transition metals. Their window is too short to settle which model a desk should run for any one metal. Their evidence supports a narrower claim: no specification wins everywhere, and a poor choice in a thin market can be expensive. This is a model-risk warning, rather than a source of trading edge.
The metals and the test
The IMF's Energy Transition Metals index supplies the universe. Data providers divide its 16 metals into 7 base, 3 precious and 6 "other" (chromium, lithium, manganese, a Chinese mixed rare-earth carbonate basket, silicon, vanadium). These markets are thin and concentrated, with little scope for substitution; volatility therefore matters to hedging and investment. The demand pressure is substantial. In the IEA net-zero path cited by the paper, 2050 demand reaches 175.4 times the 2022 level for rare earths and 16.1 times for lithium. The authors take daily LSEG prices, deflate them by US CPI and average them into 126 monthly observations from July 2012 to December 2022.
Their first step is to calculate eight features of each metal's realized variance: lag-1 autocorrelation, the sum of squared autocorrelations to lag 12, a Hurst exponent, skewness, kurtosis, the share of zero monthly returns, spectral entropy and spikiness. PCA places each metal in an instance space, a two-dimensional map of those features. PC1 chiefly captures persistence and accounts for 66.2% of the variation. PC2 chiefly captures tails and noise, adding 15.1%.
Then comes the model contest. GARCH(1,1) faces five stochastic-volatility models: log-variance following an AR(1) (SV(1)) or an AR(2) (SV(2)), plus SV(1) with jumps, Student-t errors, or a leverage correlation. Bayesian marginal likelihoods determine in-sample rank. The expanding-window exercise produces 60 one-step forecasts from January 2018 to December 2022. RMSFE and QLIKE score variance forecasts against realized variance and an adjusted squared range; weighted quantile CRPS (continuous ranked probability score) scores return densities. With realized variance as the target, SV beats GARCH in 41.3% of 160 pairwise point comparisons. The density figure is 43.8%.
Do the usual metal groups hold up?
Partly.
PC1 separates a mostly base-and-precious group from a mostly "other" group. The broad division survives on the map's dominant axis; the awkward cases lie near its boundaries. Nickel has much higher skewness, kurtosis and entropy than most base metals, alongside very low autocorrelation. Ward clustering on the eight features places lithium with palladium, platinum, silver, aluminium and zinc. Nickel and molybdenum join chromium, manganese, rare earths, silicon and vanadium; copper and cobalt cluster with lead. Nickel lands in cluster 1. For geological co-occurrence, the paper gives copper, nickel and cobalt in sulfide-rich ores as its example. The map contains only 16 points, and the authors describe it as descriptive.
SV(2) wins in-sample, then stumbles
SV(2) has the highest marginal likelihood for 11 of 16 metals, sometimes by a wide margin. Silicon scores -328.2 against GARCH's -449.9; vanadium scores -374.1 against -482.7. Silver gives GARCH its lone win, -416.8 against -418.4 for SV(1)-t.
Forecast performance is less orderly. SV(1)-L, the leverage model, delivers the largest point-loss reduction most often. Lead is the clearest success: every SV model beats GARCH at 5% on both losses.
The bad calls are large. Vanadium's QLIKE rises 345.99% under SV(1)-L, while molybdenum's RMSFE rises 183.86% under SV(1). Apply SV(1)-L to two different metals and the contrast sharpens: vanadium's QLIKE rises 345.99%, while lead's best gain is a 34.02% QLIKE cut.
SV(2) most often records the lowest density score. The authors flag an asymmetry in that comparison. GARCH densities hold maximum-likelihood parameters fixed, whereas SV densities integrate uncertainty in parameters and latent states. Across 60 forecasts, that difference can determine a tail-weighted score.
The conclusion credits GARCH with more accurate rare-earth point forecasts. RMSFE supports that reading against all five SV models. QLIKE gives the opposite ordering for SV(1), SV(2) and SV(1)-L, with gains of 15.86%, 19.11% and 12.79%. None of those three QLIKE gains reaches significance under Diebold-Mariano, even at 10%. The authors acknowledge in their next sentence that rankings depend on the loss function. My objection is more specific: the rare-earth conclusion selects the RMSFE reading without naming it.
How much can sixty months decide?
Using realized variance, Diebold-Mariano rejects equal accuracy in 20% of comparisons. With the range proxy, it rejects in 14 of 160, about 9%. That is roughly what tests at 10% yield when no effect is present. The main tables assess pairs separately; multiple-comparison control appears only in the Model Confidence Set robustness check. The paper says this check finds SV models "consistently outperform" GARCH under both point and density losses. Yet GARCH is excluded on point losses for 7 of 16 metals under squared error and 6 under QLIKE, leaving it in contention for most metals. Centre-weighted density forecasts exclude GARCH for 11 of 16. The phrase fits the density result and overstates the point-forecast result.
Rankings also move when the evaluation changes. Removing six turbulent months (54 of 60 left) raises the count of metals for which SV(2) is the best SV density specification from 7 to 10. A 66-month rolling window takes SV's point-forecast win rate from 41% to 51%.
The variance target raises a separate concern. Monthly returns use log changes in monthly average prices; the proxies use daily returns within each month. The paper says both measure the same annualized conditional variance. Monthly averaging smooths returns, so the fitted series has lower variance than the proxy used to score every model. Because QLIKE penalizes underforecasting more heavily, it may favor the model that happens to forecast high.
The authors acknowledge the sample's brevity: "Model rankings should therefore be interpreted cautiously." They argue that the turbulence and rolling-window checks show the qualitative findings are not driven solely by a few extreme months. That defence supports the absence of a universal winner. They also concede that no particular ranking survives, calling rankings sensitive to the composition of the short evaluation sample.
What a desk can use
We could not test this ourselves. Our data lacks lithium, cobalt, rare-earth and most minor-metal series. Copper and silver futures alone cannot reproduce the cross-metal comparison. We did not find a VaR or hedging P&L evaluation in the paper either; the forecasts are assessed by statistical loss.
The practical use is defensive. The authors say zero-return shares make proxies noisier for thin metals. Keep GARCH(1,1) as an anchor there, and select SV variants metal by metal. An economic test over these same 60 months, showing that per-metal selection reduces hedging error against a single default model, would change my view.