Last-window volatility does most of the work in the FESX ten-minute forecast. Thawait's continuous stress dial adds something worth measuring, though the headline R² makes that addition look larger than it is.

One coordinate for the book

Thawait uses ten-level, nanosecond-stamped snapshots of the EURO STOXX 50 front future from 2022 to 2025, leaving 987 clean days. Each 30, 45 or 60-second window supplies 15 variables from the ticks: best bid and ask, mid, quoted and relative spread, microprice, level-one sizes, trade volume, imbalances at several depths, a depth-weighted order-flow imbalance and clipped price pressure. Oracle Approximating Shrinkage pulls their covariance toward a scaled identity. The result is a 15x15 positive-definite matrix for each window; its matrix logarithm yields 120 coordinates. Thawait standardises those coordinates and takes their first principal direction as the stress dial, setting its sign with training-period VSTOXX.

The proposed shape is a connected ridge rather than a set of distinct regimes. At 30 seconds, the effective dimension is 1.4109, and the Hartigan dip on the leading coordinate is 0.00080. The author's simulations put a true continuum near 0.001 and a three-cluster alternative near 0.126. In a 17-method comparison, hard partitions have normalised mutual information, NMI, of at most 0.072 with VSTOXX tertiles. DBSCAN returns one group. On the same 8,000-state subsample, the continuous first coordinate has a 0.488 correlation with VSTOXX.

Finer partitions look better in sample. Splitting the covariance cloud into 4, 8, 16, then 27 nested groups raises eta squared from 0.0861 to 0.2626 at 30 seconds under Log-Euclidean geometry. On a single chronological half split, the out-of-sample increment instead moves from +0.0391 to -0.0813. The pooled walk-forward test deserves more weight: its Log-Euclidean gain falls from +0.02919 at the macro tier to -0.00124 at the finest, with an interval of [-0.00692, 0.00469]. Geometry complicates the verdict. Under affine-invariant geometry, the pooled gain remains positive at every tier, declining from +0.04315 to +0.01022. Only in the 2025 fold does it break: the macro estimate is +0.01338 with an interval including zero, and every finer tier turns negative. A single continuous residual coordinate adds 0.0216 [0.0118, 0.0299] over the macro dial at 30 seconds and 0.0315 at 60 seconds. Thawait's reading is that fine-scale information remains available along the continuous coordinate and gets lost when assigned to discrete groups.

Can the construction be repeated?

The text pins down much of it. Thawait gives the shrinkage weight in closed form and requires at least 17 ticks for a state. The tick filter removes spreads above 20 points and mid jumps above 50. Matrix-profile subsequences run for 30 seconds and are flagged at the daily 95th percentile. The Isolation Forest, an outlier detector that flags windows whose 40 features take few random splits to isolate, uses 200 trees, 1% contamination and seed 42 on 40 listed features. The calendar mask retains 07:30 to 15:00 UTC. It removes windows within 30 minutes of 11:45 UTC on ECB days, expiry and the prior day, and five truncated or defective dates. Forecasts use expanding-year folds, a one-day embargo and a dial refitted inside each fold. The gate keeps 80.71% of 30-second windows, a useful first checkpoint.

Input scaling remains the large open choice. On unscaled log-covariance coordinates, the leading component accounts for 83.66% to 85.27%, alongside the reported 1.4 dimension. With coordinate-wise standardisation, its share is only 26.44% to 29.31%; the paper reports this distinction and builds the dial from the standardised coordinates. The appendix says Log-Euclidean geometry responds to input units. Whether the 15 tick inputs enter as raw index points or are rescaled therefore changes the meaning of "one dimension". We did not find that choice stated for the covariance inputs. Median-and-MAD scaling is described for the window-level store.

We also did not find the decay parameter for the five-level order-flow imbalance, the price-pressure clip or the rolling reference window's length. At 45 and 60 seconds, the affine-invariant path needs a diagonal nudge of 1e-6 to 3e-6 times the mean diagonal.

Mid, bid, ask and microprice are among the state inputs. Within-window price variance therefore appears on the diagonal, and the dial may partly follow contemporaneous volatility. The persistence benchmark controls for current realised volatility, so the increment above it remains informative. The absolute R² needs to be read with that overlap in view.

The increment behind the headline

At ten minutes with 30-second covariances, support-vector regression raises out-of-sample R² from 0.5726 for persistence to 0.6821. The headline 0.7483 comes from the 60-second cell, where persistence already scores 0.6821. Roughly 0.68 of 0.75 is there before the dial enters. The author prints both figures. The increment grows toward one hour, reaching 0.0865 at 60 seconds and 0.1313 at 30 seconds, although absolute fit peaks at ten minutes. Even so, the abstract leads with 0.7483 rather than the dial's 0.0662 increment in that cell.

Across all 54 feasible cells, the best learner posts a positive increment with a day-block interval above zero. That deserves credit. A linear model captures most of the absolute fit, averaging 0.605 against 0.641 for SVR, and the paper notes that it beats persistence in every feasible cell. Nonlinear learners add 0.015 to 0.072, a substantial portion of the 0.045 to 0.131 increments over persistence.

The intraday benchmark uses single-lag persistence. We did not find an intraday HAR-style competitor. At the daily horizon, where HAR is used, the author reports no gain from the covariance state. HAR-LOB adds spread, price impact, illiquidity and order-flow toxicity; it reaches 0.8335 against HAR's 0.8298, with QLIKE 0.0760 (Diebold-Mariano p of 0.0284). It is the only model in the 90% Model Confidence Set. Adding the Log-Euclidean block to HAR-LOB lifts R² to 0.8374, while QLIKE worsens to 0.0804 from HAR-LOB's 0.0760; the Diebold-Mariano p against HAR is 0.927. The affine-invariant block falls to 0.8197 R² and QLIKE 0.1099.

The ten-minute increment weakens over time. For 30-second windows it is 0.1721 in 2023, 0.0753 in 2024 and 0.0325 in 2025. The 60-second version reaches 0.0044 in 2025. The one-hour Huber cell holds up better, at 0.1217 in 2025 with 30-second windows, so the decline depends on the horizon.

Agreement between the geometries is weakest on the 60-second grid. Only the 30-second sample satisfies both of the author's metric-agreement criteria. At 60 seconds, masking moves drift correlation from 0.1493 to -0.1666. Separately, the 45-second affine-invariant bootstrap peaks at K=4 and fails its certificate. The 2025 30-second spectral gap, 0.3103, sits just above the 0.30 cutoff.

Ten-level depth, if the desk has it

Monitoring is the practical case, chiefly because of the decile lift. The top dial decile contains 5.32 to 5.72 times the base rate of the 1% next-window volatility tail. Online change detection adds less. BOCPD records an F1 around 0.49 to 0.50, but its thresholds were swept against the same evaluation events. The paper treats those scores as comparative benchmarks rather than untouched out-of-sample tests. Quoted spread can also give the wrong impression during stress: it narrows from 0.000339 in the calmest regime to 0.000271 in the most stressed, while depth and impact worsen. Thawait tested no trading and says latency, slippage, costs and capacity would have to be addressed first.

We could not test any of this. FESX is unavailable on our platform, and our holdings are OHLCV bars without ten-level depth or quote snapshots. One-minute bars also cannot reconstruct the 30 and 45-second grids.

For me, the desk-tool test is whether the 2025 ten-minute increment clears zero against an intraday HAR-style benchmark. Table 12 reports annual point estimates only, and the 60-second cell has fallen to 0.0044.