The usable result is speed: Hybrid+Mixed needs 17 seconds per snapshot, versus 5.6 minutes for coarse-grid Cholesky. The rest is much harder to trade on.

The calculation and the data behind it

Every Deribit option settles in Bitcoin. The paper assigns the exchange 85 to 95 percent of global BTC options volume and more than $20bn of open interest. A call therefore pays max(S_T - K, 0)/S_T instead of max(S_T - K, 0).

That division drives the computational problem. An inverse put pays most on paths where S_T finishes lowest, leaving a plain Monte Carlo average hostage to a few terminal draws. Caruso uses the Mixed Estimator of McCrickerd and Pakkanen, combining Romano-Touzi conditioning with a control variate. Once the entire variance path is conditioned on, the log-return becomes Gaussian. Terminal integration then has a closed form, Black-Scholes evaluated on a conditional forward and conditional vol. Only the variance path remains stochastic. The independent driver adds no sampling noise. A timer-option payoff supplies the control variate through regression, with antithetics added. The cited reduction in variance is 5-10x from conditioning and 10-20x cumulative.

The model is rough Bergomi. Before optimisation, the forward variance curve comes from the ATM term structure. At every listed maturity, the delta-neutral ATM implied vol is squared and entered as piecewise-constant xi_0, then flat-extrapolated beyond the final tenor. Three shape parameters remain: the Hurst exponent H, vol-of-vol eta and spot-vol correlation rho.

The Bennedsen et al. Hybrid Scheme simulates the driving Volterra process at kappa = 6. It generates the fractional driver directly on the fine grid, treating the last six steps exactly and using a Riemann sum earlier. The calculation therefore avoids interpolation from a coarse grid. Three pipelines enter the comparison: coarse-grid Cholesky with log-Euler, Hybrid with log-Euler, and Hybrid with the Mixed Estimator. Differential Evolution performs the search with population 12, at most 25 generations and tolerance 0.5pp, followed by Nelder-Mead. The bounds are H in [0.01, 0.49], eta in [0.3, 4.0] and rho in [-0.99, 0].

The source files contain every listed BTC option trade on Deribit from January 2022 to December 2025, roughly 57,000 instruments. Thirty snapshots are calibrated between May 2022 and March 2025. Twenty-one cluster around seven known events, covering the day before, of and after each: LUNA/UST, FTX, SVB, spot ETF approval, the fourth halving, BTC $100k and the Trump inauguration. The outcomes were already known when those events were selected. Since the paper measures fit error and runtime rather than returns, this does not inflate the reported results.

Nine baseline dates complete the sample, three each in low, medium and high ATM IV. A snapshot includes trades within plus or minus four hours, with volume of at least 0.05 BTC, maturities from 7 to 90 days, K/S0 in [0.8, 1.2] and the OTM side only. Observations are volume-weighted.

Across the thirty snapshots, Hybrid+Mixed records mean unweighted RMSE of 22.83 implied-vol points and median 19.60. Cholesky+Euler produces 41.76 mean / 27.68 median, while Hybrid+Euler gives 39.43 / 33.67. Runtime over those same thirty snapshots is 17 s for Hybrid+Mixed, 38 s for Hybrid+Euler and 5.6 min for Cholesky+Euler.

Does 22.83 points qualify as a fit?

No. The failures contain the useful evidence.

The calm-date row needs to be read in full. The paper says all three methods fall within 1pp of one another, then reports a range of 5.12 to 14.04pp. Its 1pp statement applies only to Cholesky+Euler at 5.12pp and Hybrid+Mixed at 5.61pp. Hybrid+Euler reaches 14.04pp, nearly 9pp away. The printed range contradicts the accompanying sentence, and Cholesky is marginally best on those calm dates.

Stress accounts for the Hybrid+Mixed advantage. Its error is 22.21pp, compared with 44.90pp for Cholesky+Euler. Hybrid+Euler comes in at 38.94pp, revealing where the improvement occurs. Changing only the simulation scheme moves 44.90pp to 38.94pp. Adding the Mixed Estimator moves 38.94pp to 22.21pp. The estimator explains almost the entire gap.

On FTX day, the paper says Cholesky+Euler effectively fails to fit the surface at 124.65pp. Hybrid+Mixed returns 18.84pp. We interpret 124.65pp as a failed optimisation rather than evidence against the model itself.

Calibration error follows the vol level almost mechanically. Across all 30 snapshots, Pearson r = 0.89 and Spearman 0.91. For the nine baseline dates, mean Hybrid+Mixed RMSE rises from 5.6pp to 15.7pp and then 51.5pp across the low, medium and high ATM IV groups. The sample's best single-date fit is 3.94pp on 12 August 2023, using Hybrid+Mixed. Its worst occurs on a post-LUNA-recovery date when ATM IV exceeds 100 percent: Hybrid+Mixed reaches 88.61pp and Cholesky+Euler 102.62pp.

The conclusion acknowledges the limitation. Its three-parameter specification cannot fully reproduce the most distorted crisis surfaces, with RMSE up to about 90pp on the most extreme baseline date. A jump component or regime-dependent roughness is proposed for high-stress periods, though neither extension is implemented or tested here. The 51.5pp result in the high-IV bucket remains the best achieved by this specification. Hedging is left for future work, an honest disclosure that also leaves unanswered whether these dynamics produce a delta superior to a Black-Scholes delta.

Roughness reaches both boundaries

The abstract's roughness claim deserves the most resistance. It reports a calibrated Hurst exponent consistently near the search floor, H approximately equal to 0.01 to 0.06 in most regimes, and treats this as confirmation that Bitcoin volatility is genuinely rough. Across all 30 snapshots, Hybrid+Mixed produces mean H of 0.063. Most dates land exactly at 0.01. Average rho is about -0.29, with a substantial fraction exactly at 0.

The paper addresses those outcomes directly: the boundary solutions "are not pathological: they reflect the optimiser's preference, given a maturity-weighted three-parameter fit, for the roughest admissible dynamics whenever the short-end skew is steep, and for zero correlation whenever the smile is approximately symmetric."

Its account of the mechanics is fair, and describes a binding constraint. The residual discussion makes the same point. During FTX, negative bias gathers in the short-maturity, deep OTM put wing, where the three-parameter specification, with H fixed at its lower bound, "has exhausted its flexibility to add further roughness". Once a parameter stops moving because the model has exhausted its skew capacity, it measures the search box.

FTX Day+1 sharpens the issue. H moves from 0.010 to 0.490, and both pipelines finish together: 56.26pp for Hybrid+Mixed versus 55.96pp for Cholesky+Euler. One day after the 18.84pp fit, the advantage disappears and roughness reaches the opposite boundary.

Two further details complicate any comparison across pipelines. They return different parameters from the same surface, including mean eta of about 1.85 for Hybrid+Mixed and about 3.1 for Cholesky+Euler. Accuracy is also compared using 10,000 variance paths per maturity with n = 100 time steps. Yet the paper's own discussion calls the Hybrid Scheme's O(n^-H) strong error adequate in practice at n = 500 to 2000. Our arithmetic gives about 3 percent improvement in n^-0.01 between n = 100 and n = 2000 when H = 0.01. Any calibrated BTC Hurst exponent quoted from these three pipelines needs its pipeline label attached.

What we could not reproduce

We could not test any of this ourselves. Calibration requires Deribit trade-level inverse-option prices and volumes from a four-hour intraday window. Our options data is end-of-day and contains neither Deribit contracts nor the BTC-settled convention, preventing reconstruction of the surfaces. We have previously written about a calibrated variance model whose fit never reaches its intended trade (/articles/dspm-puts-a-volatility-clock-inside-diffusion-noise).

The result that would change my view is a rerun of the same thirty snapshots with H allowed below 0.01, followed by repricing each calibration on the next day's surface. We did not find a next-day repricing test in the paper. If lowering the floor eventually stops improving the fit, the roughness interpretation survives.