Case study · Backtest
🗓 Backtest period: 2020-01-01..2024-07-01
Spec ID: spec-calibrated-volatility-regime-defensive-allocation-us-stocks-1789597980 · Generated: 2026-09-16 23:29 UTC
Cluster: Defensive Allocation · Sub Cluster: Calibrated Vix Stress Defensive Allocation
A class-weighted ridge-logistic model estimates the probability of near-term VIX stress and continuously scales a long-only SPY allocation, leaving the balance in cash. Volatility targeting, drawdown control, and exposure smoothing further reduce risk.
Volatility stress tends to cluster, while equity drawdowns, realized volatility, liquidity, volume, and VIX state variables can contain information about elevated near-term stress risk. Investors may adjust slowly or remain structurally fully invested, allowing a systematic overlay to reduce equity exposure when estimated stress risk rises.
The strategy seeks to exchange some equity-market upside for lower volatility and drawdown. Temperature calibration converts classifier scores into probabilities, while continuous sizing avoids relying solely on a brittle binary threshold. This benefit is conditional: false alarms can leave the portfolio underinvested during rallies, and sudden shocks can occur before the predictors react.
[code omitted from public view]
| Param | Value | Notes |
|---|---|---|
| Rule type | CALIBRATED_VIX_STRESS_DEFENSIVE_ALLOCATION |
Class-weighted ridge-logistic feature-subset model with temperature calibration and continuous defensive sizing. |
| Tradeable universe | SPY&US&ETF, CASH |
SPY is the sole risky asset. |
| Signal universe | SPY&US&ETF, VIX&US&INDEX |
VIX is signal-only and cannot be held. |
| Bar / rebalance | 1d / close |
Daily MOC execution. |
| Backtest window | 2020-01-01..2024-07-01 |
New out-of-sample extension, not the paper sample. |
| Development split | 60% / 20% / 20% | Chronological training, validation, and internal untouched test before 2020-01-01. |
| Stress target | H=5, W=252, q=0.80 |
Intended event is a future five-day VIX maximum crossing the causal rolling 80th-percentile threshold. |
| Candidate predictors | 28 | Causal SPY/VIX price, volatility, drawdown, volume, liquidity, momentum, and correlation features. |
| Active predictors | 4-12 | Features are active when z_j > 0.5. |
| Ridge penalty | 10^-5 to 10^2 |
Decoded as lambda = 10^(-5+7u_lambda). |
| Classification threshold | 0.10-0.85 | Decoded as tau = 0.10+0.75u_tau. |
| Calibration temperature | 0.50-3.00 | Decoded as T = 0.50+2.50u_T. |
| Minimum model exposure | 0.00-0.60 | e_min = 0.60u_e; this can conflict with the platform cap. |
| Probability exponent | 0.50-4.00 | gamma = 0.50+3.50u_gamma. |
| Volatility target | 8%-30% annualized | sigma_star = 0.08+0.22u_sigma; uses 20-day SPY volatility. |
| Smoothing | 0.00-0.95 | kappa = 0.95u_kappa. |
| Drawdown trigger | 3%-20% | Based on 60-day SPY drawdown. |
| Initial model exposure | 100% | Operational initialization before the first valid probability; the realized position remains subject to the 50% platform cap. |
| Maximum SPY position | 50% | Platform equal-weight cap arising from the two resolved universe keys. |
| Maximum leverage | 4.0 | Platform constraint; the allocation policy itself is long-only and capped below 1x in SPY. |
| Rebalance band | 0.5% | Default minimum exposure move before emitting a trade. |
| Optimizer budget | Population 24; 40 iterations; 984 evaluations/run | Thirty independent procedure runs and thirty equal-budget random-search runs. |
| Commissions | $0.0040 per share, minimum $1.00 per order, capped at 1.00% of trade value | Charged per fill by the results module; all metrics are net of them |
| Slippage | 0 bps — not applied | Not an omission — MOC (market-on-close) fills at the auction print the backtest uses |
| Costs not modelled | short borrow fees / rebate, margin financing on leverage, market impact, exchange/regulatory/clearing pass-through fees, taxes | Excluded deliberately, not unknown |
| Execution | daily bars, MOC (market-on-close) | Exposure selected at date t applies to the next close-to-close return interval. |
| Bootstrap / alpha | 10,000 / 0.05 | Evaluation specification; expected calibration error uses 10 bins. |
| # | Concern | Status |
|---|---|---|
| 1 | Predictor timing | ✓ All SPY and VIX predictors are defined from observations dated no later than signal date t. |
| 2 | Return alignment | ✓ Date-t exposure is applied only to the subsequent SPY close-to-close return, not the return used to form the signal. |
| 3 | Close execution | ✓ Rebalances execute MOC at the official close used by the platform; VIX remains non-tradeable. |
| 4 | Training leakage | ✓ Standardization is fitted on training data and frozen; extension tuning is prohibited. |
| 5 | Forward-label boundaries | ⚠ The intended five-day labels are purged at boundaries, but the operational target fields are incomplete and the stated purge conflicts with the paper's reported effective sample count. |
| 6 | Applied costs | ✓ commissions $0.0040/share (min $1.00/order); slippage 0 bps — not applied |
H=5, W=252, q=0.80 response definition, and the operational decoding of ridge penalty, threshold, and temperature is not fully specified. In addition, purging five labels before each of two internal boundaries would reduce 1,042 observations to 1,032, conflicting with the paper's reported 625/208/209 split; the fitted classifier and paper holdout samples therefore are not fully reproducible from the specification.From 2020-01-01 through 2024-07-01, the backtest returned 22.53% with a 0.89 Sharpe ratio, 1.20 Sortino ratio, 5.63% volatility, and an -8.64% maximum drawdown. Beta versus SPY was 0.25; 511 trades produced an 80.92% win rate and 9.23 profit factor. These figures are net of the platform-applied commissions. Without matched benchmark, exposure, predictive, or calibration results, they show a profitable low-beta defensive path but do not establish superiority to SPY or reproduce the paper's model-level claims.
Recomputed here from the run's stored daily returns, because every annualized figure in the table above is derived from the LENGTH of that array (years = n / 252). Only total return and max drawdown do not depend on the row count, so when the recomputed and stored values differ by a common factor, those two are the numbers to trust.
| Check | From the run's own returns | Note |
|---|---|---|
| Daily-return rows | 1,131 | 1.00x the 1,134 trading days in 2020-01-01..2024-07-01 |
| Distinct dates | 1,131 | one row per date |
| Date span | 2020-01-02 .. 2024-07-01 | |
| Sum of daily returns | 22.53% | matches the reported total return |
| Sharpe from these rows | 0.89 | stored 0.89 |
| Volatility from these rows | 5.63% | stored 5.63% |
| Max drawdown from these rows | -8.64% | stored -8.64% |
| CAGR from these rows | 4.63% | stored 4.63% |
| Metric | Value |
|---|---|
| Total Return | 22.53% |
| Sharpe | 0.89 |
| Sortino | 1.20 |
| Calmar | 0.54 |
| Max Drawdown | -8.64% |
| Volatility | 5.63% |
| Beta vs SPY | 0.25 |
| Win Rate | 80.92% |
| Profit Factor | 9.23 |
| Total Trades | 511 |
| Symbols | 1 (SPY) |