The penalised model cuts one-day global minimum-variance portfolio variance by 8.76%, despite working with only eight assets. Whether that gain is bankable comes down to one sentence in the estimation procedure. An implementer must settle its meaning before any code can be written.
We could not test this. The paper's basket is unavailable in tradable form to us: EU allowances and Rotterdam coal have no substitute in our universe. Five of the eight series are index levels rather than tradable prices: Solactive Green Bond, Bloomberg Global Aggregate USD-hedged, MSCI Global Environment, MSCI ACWI, and the S&P GSCI 3-Month Forward Natural Gas ER. Our run would have to replace broad equities, ESG equities, bonds, oil, natural gas and gold or broad commodities with liquid US-listed ETFs. Such a run would reproduce the methodology on a substituted universe. It could neither replicate nor test the result reported here. The transition-risk interpretation, which largely explains the choice of these eight series, would also disappear.
The implementation burden is substantial. The penalised CCC-MGARCH likelihood must be coded from scratch, along with the positivity and stationarity constraints, adaptive weights and rolling penalty tuning.
Why the dense benchmark needs 136 variance parameters
Xu, Lyu and Lu begin with a constant conditional correlation MGARCH(1,1). A diagonal matrix of conditional standard deviations scales returns. The correlation matrix P comes from the correlation of standardized residuals, giving the conditional covariance H_t = D_t P D_t. This decomposition guarantees positive definiteness and motivates their model choice.
The variance recursion remains fully unrestricted: h_t = omega + A eps_squared_{t-1} + B h_{t-1}, where A and B are both N by N. At N = 8, variances alone require 2N^2 + N = 136 free parameters, alongside 28 correlations. Off-diagonal cross terms account for 112 of those 136 parameters.
The authors apply an adaptive LASSO penalty directly to the Gaussian quasi-likelihood. A preliminary unpenalised fit supplies a separate weight for every coefficient: 1 over max(|theta_tilde_j|, 0.005). Small first-pass coefficients receive a heavy penalty, while large ones receive little. The covariance forecasts then determine GMV and constrained mean-variance weights. If estimation noise dominates those 112 off-diagonal cross terms, setting them to zero should steady the weights and reduce realised portfolio variance.
The sample contains daily close-to-close log returns times 100 from January 2018 to June 2025, with 1,509 balanced observations. Its eight series comprise a green bond index, a USD-hedged global aggregate bond index, MSCI Global Environment, MSCI ACWI, front-month Brent, a GSCI three-month forward natural gas index, ICE Rotterdam coal, and EEX EU allowances. Volatility varies by a factor of nine, running from 0.90 for the hedged aggregate to 8.046 for coal. EUA alone has a significant positive drift, measured at 0.333 a day.
Out-of-sample forecasts use a rolling 1,000-observation window, re-estimated every fifth observation, at horizons of 1, 5 and 22 days. The losses are QLIK and Frobenius against RC_t = r_t r_t', together with GMV and constrained MV portfolio variances. White's (2000) reality check, using a stationary bootstrap, tests all of them.
Sparsity pays at N = 8
The paper repeatedly calls the setting "high-dimensional". A footnote acknowledges that the term describes the parameter space rather than N. The arithmetic makes the stronger case. With a 1,000-observation rolling window, the dense benchmark has roughly 7.4 observations for each variance-equation parameter.
About 37% of the off-diagonal entries in A and B are significant at 5%. The authors interpret this as a dense transmission structure partly created by sampling variation. Penalisation raises the log-likelihood from 10.2094 to 11.2304, while BIC drops from 916.4390 to 636.2674. That BIC change should be read descriptively. The paper openly uses surviving nonzero coefficients as the effective degrees of freedom, leaving penalty selection uncharged.
The retained structure has a clear asymmetry. Mass in A moves toward the diagonal instead of shrinking evenly. The AGGB own-effect rises from 0.1598 to 0.2377, and STOCK moves from 0.1723 to 0.2370. Almost every cross term in A disappears, although several channels remain in the persistence matrix B. In the coal equation, the gas shock loading survives at 0.1900 in A and 0.1430 in B.
Daily impact transmission nearly vanishes. The remaining transmission works through persistence.
Which sample builds the weights?
Two passages give conflicting accounts of step one. According to the method section, step one estimates the unrestricted CCC-MGARCH "on the full sample" and collects theta_tilde. Its expression for lambda_max evaluates the criterion at that full-sample estimate. The empirical section instead assigns the first 1,000 observations to a training subsample used "to construct the adaptive weights". One-step QLIK over the final 509 observations then selects lambda-hat. These descriptions produce different estimators, and the paper leaves too much ambiguity to identify which one was run.
The issue is unusually consequential because 1,509 minus 1,000 is 509. Suppose lambda-hat is selected once on the validation block, then held fixed during the rolling exercise. The same observations used to minimise one-step QLIK would later score the forecasts. Rebuilding the weights and lambda within each of the roughly 102 re-estimations would yield a clean exercise. It would also require 30 penalised fits of a constrained 136-parameter problem in every window, using Matlab's fmincon. The authors need to specify which procedure they used.
Two smaller implementation choices remain. Negative off-diagonals appear in both estimated A and B matrices, including an AGGB entry of -0.0672 in the penalised coal equation. Constraints on the variance recursion itself keep h_it positive, but the paper leaves enforcement inside the optimiser to the implementer. The MV target return is fixed at 10% following Engle and Kelly. Because returns are scaled by 100, that target also requires a units convention.
Mostly a QLIK result
QLIK falls by 3.40% at h = 1, 4.11% at h = 5 and 4.10% at h = 22. The corresponding Frobenius improvements are 0.53%, 0.48% and 0.80%. At h = 1, the comparison is 3.40% against 0.53%. Across the three horizons reported in Table 4, the likelihood-based metric changes five to nine times as much as the Frobenius norm. Forecast-matrix level accuracy therefore moves very little.
The proxy may explain part of the gap. RC_t is one outer product of daily returns, and the authors identify higher-frequency realised measures as future work.
GMV realised variance declines from 1.382 to 1.261 at h = 1, from 1.456 to 1.334 at h = 5, and from 1.524 to 1.402 at h = 22. Every result is significant at 5%. Constrained MV portfolio variance improves by 8.61%, 8.38% and 5.65%. These are gross portfolio variances by construction, as required by the Engle and Kelly loss function. Lower variance is the objective; requesting a Sharpe would change the question. Turnover is the missing figure I would want, since sparse and dense variance structures should diverge there most visibly.
Correlations remain fixed
The authors state this in both the model section and the conclusion. They describe a DCC version as "a straightforward extension" and time-varying correlations as "a natural next step". P is re-estimated for each 1,000-observation window, with the window advancing every fifth observation. It remains frozen within each window and throughout the complete 22-day forecast path.
The sample spans March 2020 and February 2022. Every basis point of the 8% to 8.8% GMV improvement comes from pruning cross terms in the variance equations while holding the correlation structure fixed. I would expect a DCC version to retain some of the 8.76% h = 1 gain and surrender a substantial portion. A frozen P makes the variance equations absorb movements in correlation that a DCC layer would model directly.
The dense CCC benchmark does give the experiment a clean treatment comparison. The paper asks precisely "does adaptive regularisation improve multivariate covariance forecasting relative to standard MGARCH?" Yet this is also its sole comparison. Sparse CCC against DCC remains untested, as does sparse CCC against a 60-day EWMA covariance.
Readers cannot verify two further claims. Standard errors for the penalised estimates are "available upon request", preventing assessment of the surviving spillover channels' significance from the paper alone. The table also gives a Ljung-Box(20) statistic of 10.501 for squared returns on the hedged aggregate bond index, beside a reported 5% critical value of 31.41. The surrounding text describes the statistics in every column as significant.
Section 4.4 acknowledges the basis of my objection. The paper contains no formal subsample or break tests, and the authors present the rolling window as an informal stability assessment. They propose explicit subsample splits and formal break or stability tests as useful directions for future research. Their defence rests on forecast performance being evaluated across many overlapping subsamples and changing market conditions. That argument works for forecast scoring. Its force for the penalty depends on the unresolved estimation fork.
If lambda-hat was fixed once using the final 509 observations, a single value chosen on part of the evaluation period carries through every one of the roughly 102 re-estimations. Rolling the window does not sever that link.
Confirmation that the weights and lambda were re-derived inside every rolling window would make this a genuine result about variance-equation regularisation at moderate N, and I would next ask for the turnover figures. If they were fixed on the final 509 observations, the Frobenius result is the number to trust.