A normal cutoff would make this daily covariance alarm fire too often. With five assets and a thousand intraday returns per day, Golosnoy and Seifert's preferred test has a simulated 99% cutoff of 2.8972, against 2.3263 for a standard normal. The difference matters before a minimum-variance rebalance: a practitioner can check whether today's realized covariance matrix is plausible under the fitted model, yet the advertised rejection rate depends on simulating that model. The asymptotics do not supply a usable cutoff. The authors agree and recommend simulated critical values in applications. Their values come from one parameterisation drawn 10^7 times; another fitted model calls for another simulation. The tables also contain two details worth keeping in view: some small shifts reduce rejection below the null rate, and the printed null rejection rates sit near 5%.
Every result in the paper comes from Monte Carlo draws of its own model with five assets. The paper makes no empirical claim. It uses no market dataset, so our ETF run below can neither confirm nor contradict the authors' findings.
Inside the covariance model
The authors adapt the old VEC-GARCH to realized covariance matrices. Conditional on the past, each day's matrix R_t is Wishart with v degrees of freedom, centred on a latent scale matrix S_t. Here v plays the role of the number of intraday returns. Stack the 15 distinct entries of a 5x5 matrix into a vector, the paper's vech, and the scale follows s_t = ω + A r_{t-1} + B s_{t-1}. The authors call this the realized VEC-CAW; CAW denotes the conditional autoregressive Wishart family. The unrestricted model has parameters of order O(n^4). For simulation they use a restricted CAW(1,1) that keeps S_t positive definite.
Proposition 1 supplies the new moment theory. The unconditional mean is σ = (I − A − B)^{-1}ω. It exists if and only if every eigenvalue of A + B has modulus below one. The proposition also gives a necessary and sufficient condition for a finite second moment and closed forms for the autocovariances. Those results support four statistics for a single day's r_t. Two conditional statistics measure distance from latent S_t using the Wishart covariance G_t. Two unconditional statistics measure distance from the long-run mean σ using Cov(r_t). The Mahalanobis versions whiten by the inverse covariance; the Euclidean versions leave that inversion out.
Earlier work by Hafner (2003) and Sliwa and Schmid (2005), both cited in the paper, developed moment theory and Mahalanobis monitoring for VEC-GARCH on daily returns. The change here is to realized measures. Under the Wishart assumption, S_t determines G_t completely, making the conditional statistics computable once S_t is available.
The simulation fixes five assets, hence p = 15, and sets v to 100, 400 or 1000. The authors associate those values with roughly 4-minute, 1-minute and 25-second sampling. Long-run variances are 5, covariances are 2, correlation is 0.4, and the spectral radius is 0.8865. Following a 100-day burn-in, they shift one day's intercept. The shift matrix uses diagonal δ to scale variances and off-diagonal ρ to add comovement. The null is δ = 1 and ρ = 0.
The intended use is clear: the authors point to global minimum variance construction. A rejection prompts a review of the covariance input before rebalancing. They make no return-signal claim.
Why simulate the cutoff?
The asymptotics require p → ∞ and p/v → 0. The authors acknowledge that applications can hardly satisfy these conditions; in their words, the asymptotics "mainly have the role of a theoretical benchmark." At v = 100, the simulated 99% cutoffs are 3.3380 for unconditional Mahalanobis and 4.1584 for unconditional Euclidean. Even at v = 1000, they remain 2.8972 and 3.7268. N(0,1) gives 2.3263. Apply that normal cutoff to 15 covariance elements and rejection will exceed the nominal 1% rate.
Estimation risk appears in the authors' discussion only for S_t. It reaches the critical values as well. Their 10^7 replications use one parameterisation, while another fitted model needs its own simulation and brings estimation error that the power tables do not face.
One printed detail remains unreconciled. The text says the power calculations use simulated 99% quantiles. In the variance table, however, the null rejection rate for unconditional Mahalanobis at v = 100 is 0.0502; in the correlation table, null rates hover around 0.049 to 0.051. Those rates resemble a 95% cutoff. The power figures below are reported as printed.
Equation (11) states that r_t becomes unconditionally normal as v grows, and the paper treats the unconditional Mahalanobis distance as chi-square under the null. Yet its quantile falls from 3.3380 to 2.8972 when v rises tenfold and remains above 2.3263. Approximating a chi-square with 15 degrees of freedom by a normal accounts for the gap: at that dimension, right skew pushes upper chi-square quantiles higher. The authors describe p = 15 as too small. A chi-square cutoff is therefore the natural first choice ahead of a normal cutoff for Mahalanobis. At small v, simulation still matters: the cutoff is 3.3380 at v = 100.
Where Mahalanobis loses power
For variance shifts, Mahalanobis is clearly stronger. At v = 400 and δ = 1.2, unconditional Mahalanobis power is 0.4292, versus 0.1198 for unconditional Euclidean. Cut variance to δ = 0.7 at the same v and the figures are 0.6312 versus 0.1391. At v = 100, an increase to δ = 1.6 has power of 0.8039, while a decrease to δ = 0.5 has only 0.1855. Conditional Mahalanobis does marginally better on increases, reaching 0.4338 at δ = 1.2 and v = 400. It needs latent S_t, though, and the authors favour the unconditional version because it depends only on model parameters.
Small correlation increases expose the weak spot. The authors call the Euclidean edge "rather remarkable" and suggest Euclidean tests for correlation-driven changes. With v = 100 and ρ = 0.2, unconditional Euclidean detects 0.0972 against unconditional Mahalanobis at 0.0456. At ρ = 0.3, the figures are 0.140 and 0.080. Mahalanobis regains the lead by ρ = 0.5, at 0.3652 versus 0.2721. At v = 1000 it dominates too, with 0.8432 against 0.4201 at ρ = 0.2. The 0.0456 result sits uneasily with the authors' closing description of "decent power for all considered types of changes" for unconditional Mahalanobis.
The tables contain another feature we did not see discussed in the text. At v = 100, unconditional Mahalanobis rejects less often for several shifts than it does under the null: 0.0297 at δ = 0.7, 0.0281 at δ = 0.9, and 0.0366 at ρ = 0.1. Power below size makes the test biased there. Euclidean rejection also dips, reaching 0.0444 at δ = 0.9. Under coarse sampling, a small volatility decline or a mild rise in correlation can quiet the alarm.
These experiments shift a single day. The paper examines neither persistent breaks nor repeated daily testing. At the intended 1% rate, 252 sessions would produce about 2.5 false alarms a year with a perfect model. At the roughly 5% null rate printed in the tables, the figure is about 12.6.
Our SPY, TLT and GLD run
The authors report no portfolio performance. The figures here are ours, with no result of theirs to compare them against.
We formed realized covariance matrices for SPY, TLT and GLD from synchronized one-minute closes, requiring at least 150 returns a day. Through t−1, we refit a restricted CAW(1,1) on the trailing 252 sessions. For day t, we compared the unconditional Mahalanobis statistic with a simulated 99% cutoff. A rejection moved the allocator from the CAW forecast to the trailing 252-day mean covariance, shrunk 50% toward its diagonal. The minimum-variance portfolio kept 90% invested and 10% in cash, with each ETF weight between 0 and 1/3. Trades filled at the next open; commissions were $0.004 a share.
From 2020-01-02 to 2024-07-01, the gated book returned 24.30% net of those commissions. Annualised volatility was 9.97%, Sharpe was 0.54, and maximum drawdown was −22.34%.
The constraint explains more than the gate. Three positions capped at 33.33% must together reach 90%, forcing each above 23.33%. Each weight can therefore move only between 23.33% and 33.33%. Set against 9.97% volatility, the −22.34% drawdown mostly describes holding those three ETFs over that period. Beta to SPY was 0.25, so the drawdown was not mainly an equity drawdown. We also lack a never-test comparison and cannot tell whether the gate reduced risk or turnover. Our calibration used 100,000 null paths and was refreshed every 21 decisions; the authors used 10^7 replications. Our 1% tail cutoff is consequently noisier than theirs. This single automated pass says little about the test and considerably more about our design choices.
Evidence still missing
Within the authors' model, unconditional Mahalanobis earns their recommendation. Its reported size is near 0.05, and power for a variance cut to δ = 0.7 rises from 0.6312 to 0.9975 as v goes from 400 to 1000. The remaining question is how it behaves on real realized covariances, where microstructure noise tests the Wishart assumption and estimated parameters determine the critical values. Evidence that rejection days precede larger out-of-sample covariance forecast losses would move me from seeing it as an interesting diagnostic to wanting it wired into a rebalance.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.