The S&P 500's overnight return deserves its own variance state. Over the 2000 to 2020 sample, trading outside market hours produced 0.407 of the index's 0.803 cumulative log-return while contributing 6.4% of daily return variance.
Half the drift, one fifteenth of the variance.
That asymmetry justifies separate variance states for the two sessions, which is what Mpanda's paper supplies. Whether a trading desk gets enough value from the added machinery is a harder question, and the paper's own abstract gives away part of the answer.
We did not run this experiment. Every figure below comes from Mpanda and was computed on the index itself. We would trade SPY, while the index supplies readable prices without being a holdable instrument. The mechanism carries across. The reported forecast losses cannot be transferred directly.
Inside the CouRTES filter
CouRTES, the author's name for the coupled realised real-time EGARCH, tracks two log-variance states, overnight and intraday. Each has separate persistence and leverage responses. A 2x2 matrix links the lagged states, while another connects the lagged shocks. An overnight gap can therefore lift tomorrow's intraday variance. A bad afternoon can lift tonight's overnight variance. Daily variance is the sum of the two.
Two further inputs drive the recursion. The realised measure combines the squared overnight return with 5-minute realised variance and enters as a log term. A same-day proxy then updates the filter with current-session information. The proxy divides the open-to-close return by a 20-day rolling standard deviation of intraday returns.
Filtered historical simulation sits above the filter. Standardised return residuals and measurement residuals are resampled as pairs, sent through the recursion, and converted into a predictive distribution rather than a Gaussian one.
The empirical design uses a rolling window of 1000 observations and refits every 21 origins. It contains 250 forecast origins from 30 May 2018 to 29 May 2019, covering horizons from 1 to 252 days. The benchmarks are Realised EGARCH, Coupled GARCH(1,1), realised real-time GARCH, EWMA at lambda = 0.94, and HAR-RV.
QLIKE wins, then HAR-RV catches up
QLIKE provides the paper's strongest evidence. CouRTES records the lowest loss at 11 of 14 horizons within the GARCH family. Its median improvement over Coupled GARCH(1,1) is 6.23%. The largest reductions are 25.43% at H=147 and 23.40% at H=168. Against realised real-time GARCH, the median improvement reaches 43.78%.
Diebold-Mariano tests using Newey-West errors report lower mean loss in 36 of 42 comparisons with the three GARCH-family benchmarks. Of those comparisons, 22 are significant at 5% and 26 at 10%. At the 10% level, the Model Confidence Set retains CouRTES at 13 of 14 horizons and removes it only at H=168. CouRTES also improves on EWMA at all 14 horizons, by a mean 33.47%.
The paper's own abstract concedes where the result stops: CouRTES "improves over the EWMA benchmark across all reported QLIKE horizons, but remain competitive to HAR-RV model only on realised volatility point forecasting." HAR-RV uses three regressors, the daily, 5-day and 22-day averages of realised volatility. It frequently remains in the confidence set beside the coupled model. In the MCS input table, CouRTES has the lowest mean QLIKE at three horizons, H=63, H=84 and H=105.
Coupling and the same-day channel buy plenty against the GARCH cousins. HAR-RV removes the consistent edge. Three of fourteen horizons give CouRTES the lowest mean QLIKE, and the paper's own DM tests do not establish consistent superiority over HAR-RV. For a variance estimate used to scale SPY exposure, the 2009-vintage answer already does the job.
The remaining case therefore rests on the distributional layer, where the comparison set disappears. Table 1 assigns FHS to CouRTES alone, leaving its coverage result unopposed. Average empirical coverage is 0.878 against a nominal 0.90, and the range runs from 0.752 to 0.948. Mean interval width stays near 1.90e-4. The calibration looks plausible, and stable width indicates that wider bands did not purchase the result.
No benchmark received the same treatment. We consequently cannot tell whether adding an FHS layer to HAR-RV would produce better results. A risk manager needs that comparison, and the paper does not run it. There is no VaR or ES backtest. The paper also contains no trading-cost or P&L evaluation. Risk management and volatility-index construction supply the stated motivation, while VaR and ES appear only as a future extension.
Near a year, the evidence thins out
Coverage falls to 0.768 at H=231 and 0.752 at H=252. QLIKE worsens beyond H=189. Mpanda attributes the decline to errors in own-segment persistence and cross-segment transmission, which accumulate through the recursion before aggregation.
A simpler candidate deserves attention. With 250 origins and H=252, the long-horizon loss paths overlap almost completely. The result is effectively one observation carrying 250 labels. The paper applies HAC and Newey-West errors to handle serial dependence in multi-step loss differentials. HAC cannot recover information from a loss series that is effectively one observation, leaving those horizons with almost no test power.
The calendar compounds the problem. The effective sample ends 3 June 2020. By our arithmetic, 252 trading days after the final origin on 29 May 2019 falls in roughly that week. The last H=252 targets therefore reach the spring of 2020, including the realised-measure maximum of 9.998e-3 on 16 March 2020. Those H=252 losses amount to a handful of overlapping paths through one episode.
Treat H up to about 100 as evidence. Anything past 189 is a single draw.
Who knows the same-day proxy?
Proposition 2.1 says the segment variances are measurable with respect to information through day t-1. Yet the paper defines that information set as the sigma-field generated by returns and realised measures up to t-1. The same-day proxy uses the day-t open-to-close return, so the intraday variance state is updated with the return whose variance it describes. In the proof, the paper lists the proxy among the day t-1 measurable objects. We do not think that follows from the stated definition.
The issue matters more in forecasting than in filtering. Producing a one-step-ahead variance requires the next session's intraday return, standardised by a rolling window. The FHS description explains how the standardised return residual and measurement residual are resampled as pairs. It does not explain how psi(z-tilde) is propagated at h greater than or equal to 1.
Anyone rebuilding the model would have to guess. An ambiguous estimation step similarly blocked reproduction of the Adaptive LASSO-MGARCH paper (/articles/adaptive-lasso-mgarch-tested-against-a-dense-benchmark-only). We also did not find fitted coefficient values, so the cross-segment spillover parameters motivating the design never appear.
One more issue matters before use. Daily variance equals the sum of the two segment variances, with no cross term. Measured correlation between overnight and intraday returns is about 0.25, and the 2Cov term represents 11.5% of Var(r_t). This treatment is consistent with the paper's own target because its realised measure is additive as well. Applied to a close-to-close VaR limit, the omission biases daily risk downward.
The overnight and intraday tails make the underlying idea worth keeping. Overnight skewness is -3.30 and excess kurtosis is 108.84, versus -0.20 and 8.42 intraday. Those distributions differ sharply, while the 11 of 14 QLIKE wins over single-state GARCH benchmarks indicate that a single daily filter leaves information unused. I would use the two-state structure. Before paying for the FHS interval layer, I would want it tested against HAR-RV using identical residuals.