Nyblom statistics of 6.6249 to 15.9059 make drifting ω the result that matters if a GARCH filter sets your rates risk. The regime account layered over it has much thinner support.
Balcerek and Wronka study daily log changes in the USD at-the-money forward 1Y×10Y swap rate. Their sample runs from 1 January 2007 to 29 September 2023 and contains roughly 4,365 observations. The series was purchased from S&P Global, and the paper says it "has been calibrated and processed based on internal expertise and proprietary models." It spans the LIBOR, OIS and SOFR discounting eras without a break.
They estimate three models:
- GARCH(1,1).
- GJR-GARCH(1,1), where a γ term adds variance after negative shocks.
- A two-regime Markov-switching GARCH (MSGARCH). A hidden Markov chain shifts the process between latent states, each with separate (ω, α, β) parameters.
The Nyblom test asks whether each parameter's likelihood score changes through the sample. It rejects constant ω in all four single-regime specifications, using a 1% critical value of 0.748. By contrast, the α, β and γ statistics are all at or below 0.1692. The short-run coefficients therefore appear stable while long-run variance, ω/(1−α−β), moves. Yet the Nyblom section identifies swaption-vol returns as the input, leaving unclear which series generated the reported 6.6249 to 15.9059 range.
MSGARCH supplies the proposed fix. Its regime parameters remain fixed, while changing filtered state probabilities absorb the drift. The authors describe that benefit as structural and acknowledge that MSGARCH "does not materially outperform single-regime models in short-horizon tracking."
We could not rerun any of this. The swap fair rates and swaption vols are proprietary S&P Global series. Substituting Treasury futures would examine volatility-based sizing in another market, so none of the estimates below would transfer. The discussion that follows concerns the choices an implementer would have to make.
The signs in the likelihood table
The reported log-likelihoods are −10708, −10718 and −10722 for GARCH, GJR and MSGARCH. Alongside them, the paper gives AIC values of −21407, −21426 and −21427. Those signs cannot coexist under AIC = −2logL + 2k: a negative log-likelihood does not yield a negative AIC. The AIC and BIC columns can be reproduced by reading the likelihoods as +10708, +10718 and +10722, with T near 4,365. For MSGARCH, −21444 + 9·ln(4365) gives about −21369, matching the table. Dropping the apparent typographical minus signs also leaves MSGARCH with the highest likelihood, in line with the paper's claim.
A separate inconsistency remains. The likelihood of 10722 conflicts with the −10771.29 shown in the Student-t parameter table, or 10771.29 if the same sign error applies. The paper may contain two MSGARCH fits, or one fit with two reported likelihoods. We could not determine which. In either case, MSGARCH leads GJR by 1 point on AIC, while BIC favours GJR by 25, as the abstract concedes.
One AIC point bought with four extra parameters decides nothing.
The innovation distribution also requires a judgment call. Adjusted Pearson χ² values are 102.17 for the Normal, 4364.00 for the Student-t, 133.87 for GED and 36.56 for Johnson SU. Even so, Student-t remains the baseline MSGARCH specification, with ν = 14.59. The Johnson SU MSGARCH likelihood of −9398.62 belongs to standardized residuals from a GJR fit on rescaled returns. It is therefore incomparable with −10771.29. Although the paper describes the Johnson SU improvement as "marginal," we found no like-for-like figure supporting that description. It also calls the Student-t shape statistics of 0.2919 and 0.2895 "strong instability," despite both falling below the 10% critical value of 0.353.
Which state is the crisis?
Under the Student-t fit, regime 1 has ω = 9.72×10⁻⁵, α = 0.352, β = 0.599 and p11 = 0.450. For regime 2, ω = 4.11×10⁻⁷, α = 0.033, β = 0.958 and p22 = 0.921. The paper labels regime 2 as the high-volatility state and associates it with 2008 and COVID.
The parameter arithmetic suggests the reverse labeling. Regime 1 has implied unconditional variance of about 9.72×10⁻⁵/0.049, equal to roughly 2.0×10⁻³. Regime 2 gives about 4.11×10⁻⁷/0.009, or 4.6×10⁻⁵. That figure is roughly forty times lower. Expected durations calculated as 1/(1−p) are 1.8 days for regime 1 and 12.6 days for regime 2. The paper says the transition matrix "indicates extended durations within each regime." Yet regime 1 behaves more like a brief spike state, while a filtered probability held near one from 2007 to 2009 is hard to reconcile with a 12.6-day expected stay.
The filtered-probability chart carries the event interpretation. In the collateralisation discussion, the authors say β falls and α rises between the LIBOR and OIS eras. We found no subsample estimates supporting those statements. A refit should assign state labels from implied variance rather than β.
Coverage without formal tests
Across 1500 test days, one-day 95% interval coverage is 94.00% for GJR, 93.80% for GARCH and 93.8% for MSGARCH. The corresponding results over 1000 days are 93.30%, 93.20% and 93.0%. At 1000 days, the binomial standard error is about 0.69 points, placing 93.0% nearly three errors below nominal coverage. Breaches reach 7% against the promised 5%, which is 40% more than advertised. This 1000-day window produces the weakest result.
Shorter windows come closer to nominal coverage or exceed it. At 500 days, the figures are 94.40%, 94.60% and 95.0%. At 300 days, they are 95.67%, 95.33% and 96.0%.
The intervals use mean ± 1.96σ, a normal quantile, even though Student-t innovations are fitted for MSGARCH. We did not find Kupiec or Christoffersen tests, or a loss comparison using realised or implied vol. Model treatment also differs. According to the paper, "we estimate MSGARCH once on the full sample and apply a Hamilton-style filter," whereas GARCH and GJR are re-estimated day by day. The authors regard rolling MSGARCH re-estimation as conceptually misleading. Full-sample estimation nevertheless allows test-period information into the reported coverage. Across the five windows, MSGARCH is closer to 95% than GJR only once (500 days: 95.0% vs 94.40%).
One rebuild tracks SRVIX
The benchmark comes from the authors' reconstruction using vendor swaption vols, including one SABR-priced version. They never use the published CBOE index. Correlation with the GARCH filter is −14% when the rebuild uses 1000 SABR-priced swaptions, and 1.5% with the 11-swaption CBOE formula. Only the Carr-Madan fair-volatility-strike construction is described as following the GARCH filter closely, yet no correlation is supplied. Its "little adjustment for annuity" also goes unexplained. The reported −0.72 correlation between the rate and ATM vol is based on lognormal vol, which scales inversely with the rate level and therefore contains a mechanical component.
The ω finding remains the paper's strongest result, subject to clarification of which series entered the Nyblom test. A rolling check of long-run variance follows naturally from that evidence. The argument for spending nine parameters on MSGARCH would become persuasive with one result: rolling MSGARCH re-estimation whose formal coverage tests beat GJR. The authors decline that test because they consider it conceptually misleading.