A four-parameter ARMA(1,1) on log realized variance keeps pace with a fractional Ornstein-Uhlenbeck process out of sample at the one-day horizon. Against the HAR benchmark, average QLIKE across ten ETFs is 0.7942 for the ARMA, 0.7870 for fOU and 0.7849 for fractional Brownian motion. The authors make the comparison plainly. Their models' accuracy, they write, "is comparable to that of the rough continuous-time models but much easier to estimate by standard off-the-shelf software." The claim rests on statistical support that appears under QLIKE and disappears under MSE.

That makes Bennedsen, Christensen, Christensen, Yu and Zhang worth rebuilding.

Fractional noise reduced to one lag

Daily samples of fBm and fOU are approximately an AR(1) driven by fractional Gaussian noise. So approximating a rough volatility model comes down to approximating fGn. The authors numerically recover the Wold coefficients of fGn, the moving-average weights that map white noise into the series, from its spectral density. That density has a closed form in the Hurwitz zeta function. They apply a Szegő factorization, then a cepstral recursion, and inspect the coefficient decay.

The decay is quick. At H = 0.2, the first coefficient is -0.4588 and the second is -0.0788. Across H in [0.05, 0.5], roughly 75% to 100% of the lagged variance in the MA representation comes from the first coefficient, with the largest share in the rough range. As the paper puts it, "The first lag therefore does almost all of the work." c1 also moves almost linearly with H: -0.6568 at H = 0.1, -0.4588 at 0.2, -0.2926 at 0.3 and -0.1419 at 0.4. The authors describe that linearity as surprising. It is. The ARFIMA fractional differencing parameter behaves very differently, piling up at its -0.5 boundary once H falls below about 0.2.

Truncating the Wold expansion at lag one produces an MA(1), which the authors attach either to an AR(1) or to a HAR cascade. Exact maximum likelihood estimation runs through a Kalman filter. They call the resulting specifications the "rough" AR model, an ARMA(1,1), and the "rough" HAR model, HAR with MA(1) errors. Both use log RV.

Rough paths repeatedly reverse over short intervals. In discrete time, the signature is a steep first-lag drop in the autocovariance. For fGn, that drop equals 2^(2H-1) - 1, about -0.34 at H = 0.2. An AR(1) fitted to data that prefers ARMA(1,1) drags the persistence estimate toward the first autocorrelation. Table 3 shows the effect. Plain AR rho ranges from 0.7854 to 0.8546 across the ten ETFs. Rough AR rho lies between 0.9328 and 0.9616, while theta runs from -0.4076 (XLF) to -0.4800 (XLU). The MA term takes the short-run reversal, allowing the autoregressive term to reflect the persistence in the data.

The sample contains ten ETFs: SPY and the nine sector funds partitioning the S&P 500. The authors construct five-minute RV from TAQ over January 2010 to December 2024, giving roughly 3,770 days. Out-of-sample forecasts use a 500-day rolling window. It advances one day at a time, with every model re-estimated after each move. Forecast horizons are one day, one week and one month. Competitors include HAR, HARQ, HARJ, HARS, log-AR, log-HAR, fBm and fOU. The broader check uses a 40-stock panel from the VOLARE archive, spanning 2015-01-02 to 2026-01-30.

Our data cannot reproduce theirs

We have not tested their method. Their realized variance is constructed from TAQ tick data under sampling conventions we cannot match. We have 1-minute OHLCV bars, and five-minute RV assembled from those bars is a different series. A rebuild also requires choices on the intraday sampling frequency and any noise correction.

Every model sends its forecasts through an "insanity filter". Forecasts outside the estimation window's RV range are replaced with the window mean. Log forecasts return to levels through exp(y_hat + s²/2), where the error variance is model-implied. A correctly specified log model benefits from that correction.

One-day gains, monthly decay

At one day, average MSE relative to HAR is 0.7888 for log-RAR, 0.8051 for log-RHAR, 0.8044 for fBm, 0.8223 for fOU, 0.8374 for log-HAR and 0.9372 for log-AR across the ten ETFs. The rough specifications hold a similar advantage at one week, with log-RAR MSE at 0.7728 and log-RHAR at 0.7773. QLIKE gives the same broad ordering at one day: log-RHAR records 0.7925 and log-RAR 0.7942.

MSE reverses the ranking at one month. Plain log-HAR leads the ten-ETF average at 0.9813, compared with 0.9968 for log-RHAR and 1.0185 for log-RAR. fBm deteriorates to 1.0752. The authors present this as an exception and say the model differences there are minor. Most of the 40 stocks show the same reversal. Under MSE, every model remains in the model confidence set at every horizon. The resulting "win" is therefore an average-loss ranking without statistical separation. QLIKE at one month favors fOU at 0.8939. Log-RHAR follows at 0.9110 and log-RAR at 0.9250, both ahead of log-HAR's 0.9374. Under squared-error loss, the MA(1) term behaves like a short-horizon adjustment. A first-lag correction should have little influence on a 22-day average.

The statistical case is one-sided by loss function, and the paper reports both sides. At one day under QLIKE, log-RAR, log-RHAR, fBm and fOU are never removed from the model confidence set. Most ETFs exclude the HAR-type models, with SPY HAR p = 0.0120. MSE separates none of them. The smallest MSE p-value in the Appendix B table is 0.1622, leaving every model inside the 90% set for every ETF and horizon. Log-RAR's 21% average one-day MSE improvement, a ratio of 0.7888 across the ten ETFs, remains statistically indistinguishable from the benchmark under the authors' own test. QLIKE carries the argument. Its heaviest penalty falls on under-forecasting variance, the error that matters to a volatility-targeted book.

The paper does not examine whether log-RHAR's 21% average one-day QLIKE reduction, ratio 0.7925, changes P&L. It contains no strategy, turnover, cost assumption or Sharpe. This is a forecast-loss exercise throughout. For short-dated options, a one-day variance error becomes hedging error. And nine of the ten ETFs are S&P 500 sector slices, making "nearly every asset" close to one market factor viewed ten ways. Table 8's 40 stocks provide a wider cross-section, all beginning on 2015-01-02.

Roughness or measurement noise?

Observed log RV can acquire a negative MA(1) through a more familiar channel. RV measures latent integrated variance with error. With AR(1) latent log variance and serially independent error, the observed process is ARMA(1,1) with negative theta by construction. Hansen and Lunde built an instrumental-variable estimator around exactly this structure.

The authors answer in two ways. HARQ is designed to exploit measurement error, yet its average one-day MSE ratio is 1.0288 and its one-day QLIKE is 1.0304, both behind the HAR benchmark. HARJ at 1.0311 and HARS at 1.0461 also fail to improve one-day MSE. If noise explained the entire result, the specification designed around noise should have recovered part of the gain.

They also impose the measurement-error variance instead of estimating it. The fixed value is 2/m = 1/39, corresponding to 78 five-minute returns during a 6.5-hour day. Theorem 2.1 of Fukasawa, Takabatake and Westphal (2022) supplies the basis: errors are asymptotically independent across days and normally distributed with variance 2/m. Re-estimating the rough AR with that noise imposed leaves theta near -0.4 for every ETF, only slightly attenuated from Table 3. The authors conclude: "Measurement error of the magnitude implied by the asymptotic theory therefore explains only part of the estimated MA coefficients." Imposing a larger variance would weaken theta further.

The distinction has little consequence for forecasting. It has to be resolved before anyone uses theta to infer a Hurst parameter.

Where the MA(1) shortcut bends

The rough AR is the place to read the implied Hurst, never the rough HAR. The Monte Carlo results leave little doubt. For fBm at n = 4,000, rough AR theta stays within 0.013 of c1(H). At H = 0.2, its average estimate is -0.467 against a target of -0.459, implying H = 0.196. Under the same setting, rough HAR produces theta = -0.278 and implied H = 0.309 when the truth is 0.2. Weekly and monthly averages in the cascade absorb short-run dependence that belongs in the MA term. The authors give the same instruction: infer the Hurst parameter from the rough AR model.

Rough HAR theta is also weakly identified in the simulations, especially for small H at n = 500, as the authors acknowledge. At H = 0.1 and n = 500, the Monte Carlo standard deviation is 0.258 around a point estimate of -0.390. A window of that length provides close to no information. The explanation is structural. Under an fOU data-generating process, the additional HAR lags are redundant once the ARMA(1,1) structure is present, leaving the likelihood flat across a broad region.

The paper discloses how boundary cases are handled. Estimates reaching the invertibility boundary are classified as spurious and rerun from interior starting values. When no interior optimum exists, the estimate is capped just inside the boundary. Empirical rough HAR thetas are more stable. On the full ~3,770-day sample, they range from -0.0817 (XLE) to -0.2844 (XLU), with standard errors between 0.0425 and 0.0469. Those fitted coefficients are precise enough. Converting them into a Hurst parameter is where the method breaks.

MA(1) remains a modeling choice. The authors explicitly decline to claim it is the best approximation, and they say MA(2) and MA(3) are straightforward to implement. The footnote reporting a BIC preference for MA(1) covers SPY log RV only.

A desk can use the parameter shift immediately. Fitting AR(1) to log RV when ARMA(1,1) better describes the series pulls persistence from roughly 0.94 to roughly 0.82. An AR(1) estimate used to set mean-reversion speed in a vol-targeting overlay or a variance swap mark is then biased. The fix takes one line of code and requires no belief in rough volatility.

Off-the-shelf software estimates the model easily. The unresolved issue lies in the realized-variance series supplied to it. The paper uses five-minute RV throughout, and we found no check using RV built at another frequency. If a one-day QLIKE advantage of this size disappears under another sampling scheme, the edge belongs to the estimator's noise.