AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

EVaR risk parity barely separates from CVaR once models match

Two of eighteen matched comparisons clear p<0.10, both in momentum deciles, and both die at 25 basis points.

2026-09-11 · 8 min read · US ETFs

Reviewing: Entropic Value-at-Risk parity for tempered stable returns · Jaehyung Choi · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Monthly Tempered-Stable EVaR-Deviation ERC with Contribution-Shock Overlay

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper

  • The exact multivariate normal tempered-stable and ICA-tempered-stable calibration, including numerical likelihood fitting and constrained EVaR-ERC optimization, must be implemented in code and may be computationally unstable in short rolling windows. A practical implementation should use robust numerical optimization, rolling-window parameter regularization, and fallback Gaussian or empirical-EVaR estimates when tempered-stable fits fail; this tests the implemented estimator rather than necessarily reproducing the paper's exact calibration.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (10)

  • deviation left undescribed by the audit (invalidates: All paper-reported universe-specific Sharpe differences, volatility reductions, turnover values, drawdowns, p-values, and claims covering the nine market-specification comparisons)
  • deviation left undescribed by the audit (invalidates: All full-sample performance levels, Sharpe differences, transaction-cost significance results, regime results, and bootstrap p-values reported by the paper)
  • deviation left undescribed by the audit (invalidates: Exact equality with the paper's portfolio weights and all associated performance, turnover, and contribution results when either added constraint binds)
  • Contribution-shock overlay: A twelve-month contribution history, median-relative shock statistic, reciprocal haircut, and post-ERC reweighting are added as a separately reported strategy variant. (invalidates: All paper-reported portfolio weights, contribution equalities, Sharpe ratios, turnover, drawdowns, and p-values as predictions for the post-overlay variant.)

6 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 14.3%Jul 2024
Sharpe
0.37
Total Return
14.3%
Max Drawdown
-22.0%
CAGR
3.0%
Volatility
8.6%
Beta vs SPY
0.28
Trades
385

What the paper reports for its own strategy

  • SECTOR EVaR-ERC ICA+NTS: cumulative return 853.79%, Sharpe 0.588, max drawdown 48.56%, turnover 6.93% (out-of-sample Jan 2000 - Mar 2026, gross)
  • SECTOR EVaR-ERC ICA+CTS: cumulative return 821.40%, Sharpe 0.579, MDD 48.39%, turnover 9.49% (gross)
  • SECTOR EVaR-ERC MNTS: cumulative return 725.76%, Sharpe 0.559, MDD 50.29%, turnover 1.97% (gross)
  • MOM10 EVaR-ERC ICA+CTS: cumulative return 1897.58%, Sharpe 0.600, MDD 58.76%, turnover 8.09% (out-of-sample Jan 1996 - Mar 2026, gross)
  • MOM10 EVaR-ERC ICA+NTS: cumulative return 1840.93%, Sharpe 0.595, MDD 58.53%, turnover 5.27% (gross)
  • MOM10 EVaR-IRP ICA+CTS: cumulative return 1815.86%, Sharpe 0.593, MDD 58.85%, turnover 8.52% (gross)

Keep the return model and portfolio rule fixed, then swap only the tail-risk measure. Entropic Value-at-Risk delivers almost nothing over CVaR. The matched comparison carries the argument. Two of eighteen matched EVaR-versus-CVaR tests reach p below 0.10, both within the same universe. Significance disappears for both after charging 25 basis points. We could not reproduce Choi's tempered-stable calibration or his constrained EVaR-ERC solve, so our own run at the end uses a substitute estimator.

How the portfolios are built

In Ahmadi-Javid's formulation, EVaR takes the infimum over a positive exponential tilt of (log MGF of the loss minus ln eta) divided by that tilt. The measure is coherent and depends on the full moment-generating function rather than a quantile. With a tractable cumulant-generating function, calculation becomes a one-dimensional minimisation over the tilt. Tempered stable models provide one.

The paper develops two multivariate constructions. A multivariate normal tempered stable projection (MNTS) keeps a weighted sum within NTS and aggregates its parameters explicitly. The other construction factorizes independent components with NTS or CTS components. Its portfolio CGF sums across components after evaluating them at the ICA-projected weights. In both cases, the tilt has a weight-dependent admissible domain inherited from the author's own earlier EVaR optimization paper.

Choi then derives the Euler contribution of each asset. Under MNTS, it depends on the time-change loading, diffusion covariance and a tilted factor. Under ICA, the mixing matrix multiplies exponentially tilted component means. The corresponding CVaR construction instead uses component means conditional on the portfolio lower-tail event, which creates the structural difference between the measures.

A location-separation result also appears: raw EVaR equals minus the fitted mean exposure plus the EVaR deviation. The Gaussian corollary deserves attention. With Gaussian returns, EVaR deviation equals sqrt(-2 ln eta) times portfolio volatility. That constant cancels from inverse-risk normalization and the equal-contribution condition, leaving EVaR-deviation IRP and ERC exactly equal to conventional volatility IRP and ERC. Without non-Gaussianity, the machinery changes no weights.

The empirical portfolios are long-only and fully invested, with monthly rebalancing, eta = 0.05 and twelve months of daily returns used for each fit. Risk contributions come from central finite differences while fitted parameters remain fixed. The study covers three universes: seven cross-asset ETFs (VTI, EFA, VWO, VNQ, AGG, TIP, GLD) out-of-sample from April 2006 to March 2026; ten value-weighted 12-2 momentum deciles from the French library, January 1996 onward; and the SPDR sector ETFs, nine to eleven names, from January 2000. Sharpe differences use a studentized circular block bootstrap, B = 4,999, with block length floor(T^(1/3)). The paper explicitly describes these tests as pairwise and unadjusted for multiple testing, an appropriate disclosure for tables with this many cells.

EVaR-ERC outperforms equal weight in all nine market-specification comparisons. EVaR-IRP does so in eight of nine. The abstract reports only the ERC half.

Equal weight confounds the result

The equal-weight comparison changes the risk measure and construction rule simultaneously. Its positive Sharpe differences therefore cannot be assigned to EVaR. Choi makes the same admission in Section 4.2: these comparisons do not isolate the effect of EVaR because both portfolio construction and the risk measure differ from the equal-weight benchmark. Yet the abstract promotes the equal-weight finding.

The largest differences show what is happening. For the seven-ETF portfolio, MNTS ERC exceeds equal weight by 0.246 in Sharpe with p = 0.011, while MNTS IRP leads by 0.177 with p = 0.016. In those same rows, CAGR is 1.52 and 1.50 percentage points below equal weight. Annualized volatility falls by 6.37 and 5.68 points. The Sharpe improvement comes entirely from de-risking a portfolio containing AGG and TIP. Any sensible inverse-risk rule, including volatility, would do the same.

The paper's cleaner control is the Gaussian EVaR benchmark. It preserves the rule while changing the distribution. A significant gross improvement appears only for momentum-decile ERC: ICA+NTS +0.014 (p = 0.041) and ICA+CTS +0.019 (p = 0.057) relative to Gaussian ERC's 0.581. Once 25 basis points are charged, the gains shrink to +0.008 (p = 0.218) and +0.009 (p = 0.370). Direct MNTS ERC in the momentum deciles remains 0.007 below the Gaussian benchmark, gross and net, with p = 0.073 and 0.063. After 25 basis points, this MNTS ERC shortfall is the sole Gaussian-benchmark comparison still below p = 0.10.

Hold the model fixed

Compare matched EVaR and CVaR portfolios using the same fitted distribution and construction. Direct MNTS produces differences close to zero across the board: -0.001 in SECTOR ERC, -0.003 in MOM10 ERC and +0.006 in XASSET ERC, with p-values of 0.642, 0.179 and 0.651.

ICA-based ERC produces larger differences, though their signs vary by universe. In momentum deciles, ICA+NTS records 0.595 versus 0.580 (+0.015, p = 0.028), and ICA+CTS records 0.600 versus 0.582 (+0.018, p = 0.065). Sector ETFs show +0.028 and +0.017, neither below p = 0.10. Cross-asset results reverse direction at -0.061 and -0.126, also with neither below p = 0.10.

Choi concludes: "These findings do not indicate a uniform performance advantage of EVaR over CVaR and should be interpreted within the investment universes, sample periods, and backtest design considered here." The qualification works in both directions. Three universes containing 7 to 11 names, rebalanced monthly from twelve-month fits, also produced the two favorable cells.

The cross-asset sign reversal contains the detail a trader should care about. ICA+CTS EVaR-IRP suffered a 47.11% maximum drawdown, compared with 23.15% for matched CVaR-IRP. The ERC pair came in at 34.56% versus 20.93%. The IRP portfolio more than doubles the worst peak-to-trough loss of the measure it would replace, while the ERC portfolio runs about 65% higher. In that universe and sample, the result runs against the purpose of buying a tail-risk measure.

Turnover sends the bill

ICA-based EVaR-ERC trades between two and three times as much as matched CVaR-ERC. One-way turnover per rebalance is 6.93% versus 2.93% for sectors, 5.27% versus 1.96% for momentum deciles, and 8.52% versus 3.98% cross-asset. ICA+CTS is more expensive again, reaching 12.76% in the XASSET ERC ICA+CTS row.

The cost grid ends at 25 basis points. For the momentum deciles, that grid does not address the main trading burden because the portfolios come from the French library, and costs from turnover inside the underlying deciles are absent from the figures. At 25bp, no matched EVaR-CVaR Sharpe difference in the paper has p below 0.10.

Section 4.5 removes the fitted location term and finds virtually no change. Comparing raw EVaR-ERC with EVaR-deviation ERC yields p-values from 0.151 to 0.993. The largest movement is XASSET ICA+CTS, from 0.678 to 0.641. EVaR is translation equivariant, while the location term enters every contribution at first order. Even so, the two constructions differ by no more than 0.037 in Sharpe. We have previously covered tempered-stable machinery that is derived carefully and then cancels from the quantity it was meant to shape (the volatility clock the denoiser cancels out).

Missing implementation choices

The backtest protocol provides enough detail to copy its broad design: eta, window, rebalance frequency, cost grid, bootstrap settings and universe composition, including the October 2015 and June 2018 additions of XLRE and XLC. We did not find the tempered-stable estimation procedure, finite-difference step h, ERC solver and convergence tolerance, or ICA component count and whitening convention.

Those choices affect the result. Contributions are numerical derivatives of a function whose admissible tilt domain changes with portfolio weights. Each model is fitted on twelve months of daily observations across shape, skew, scale, location and correlation parameters, plus a mixing matrix. We also did not find an out-of-sample row for conventional volatility ERC. Corollary 1 establishes that the Gaussian EVaR-deviation version would be exactly that portfolio, whereas Table 3 uses the raw-EVaR Normal benchmark, which retains the location term.

Our substitute run

Both the calibration and constrained EVaR-ERC solve must be built from scratch, and each can become numerically unstable on twelve-month windows. Our implementation follows its own fitting conventions: daily log returns, ECDF-MSE tempered-stable fits and FastICA with whitening. We added a contribution-shock overlay that reduces assets when their modeled contribution jumps, h_i = 1/(1+s_i). This overlay does not appear in the paper and breaks exact equal risk contribution by construction.

We tested only the seven-ETF cross-asset portfolio and only the EVaR-deviation ERC variant. The exercise tests our estimator; it does not reproduce Choi's calibration.

Our run covers January 2020 to July 2024. Net of per-share commissions with a $1.00 order minimum and across 385 fills, it produced a Sharpe of 0.37, total return of 14.32% and maximum drawdown of -21.97%. Annualised volatility reached 8.61%, explaining how 14.32% over four and a half years results in a Sharpe of 0.37. Beta to SPY was 0.28.

The paper reports gross XASSET EVaR-ERC Sharpes of 0.824 (MNTS), 0.748 (ICA+NTS) and 0.678 (ICA+CTS) over April 2006 to March 2026. The variant, window and fitting conventions all differ, so the roughly two-to-one Sharpe gap describes our run and its sample. Drawdown is more comparable. Our -21.97% falls within the paper's XASSET ERC range, from 20.57% (MNTS) to 34.56% (ICA+CTS).

The window dominates the result. Our 4.5 years include March 2020 and the 2022 joint drawdown across bonds, REITs, international and EM equity. A bond-heavy long-only risk-parity allocation loses much of its diversification there. The 0.28 beta to SPY confirms how little equity exposure the portfolio carries. The paper's twenty-year XASSET sample includes the post-2008 bond bull; ours does not.

The overlay moves in the same direction. Its reciprocal haircut on rising-contribution assets reacts to losses by cutting risk, then adds it back after recoveries, exactly the pattern whipsaw regimes punish. Commissions add a one-directional drag, small over 385 trades yet real. Sampling error around a Sharpe estimated from 4.5 years of monthly rebalances in a seven-name portfolio is also wide enough to account for plenty of the remaining gap.

The evidence that would persuade me

The paper's lasting contribution lies in its analytics: the MNTS and ICA Euler contributions, along with the clean result that EVaR risk budgeting reduces exactly to volatility risk budgeting without non-Gaussianity.

Its empirical case for EVaR over CVaR rests on two cells out of eighteen, from one universe and before costs. A single matched test could resolve the question. These universes contain 7 to 11 names, leaving tail asymmetry little room to alter a small cross-section, so the next test should use fifty names or more. The cost grid reaches only 25 basis points, where both significant cells have already disappeared, so a higher charge is needed. Every universe reports all three specifications, and the winner changes across universes. With no out-of-sample rule for choosing among them, the specification should be fixed first and only that result reported.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each calendar month-end close:
1. Select VTI, EFA, VWO, VNQ, AGG, TIP, and GLD only when each has a complete preceding 12-month adjusted-close history.
2. Compute daily log returns and fit the selected one-period daily tempered-stable model:
   a. direct MNTS; or
   b. FastICA followed by NTS or CTS component fits.
3. Center the fitted distribution and calculate portfolio EVaR deviation at eta = 0.05 by minimizing the CGF objective over the weight-specific admissible MGF domain.
4. Estimate each marginal derivative with central differences, rebuilding the admissible domain for every w ± h·e_i perturbation.
5. Set DEVaRC_i = w_i · partial_i(Dc) and solve on the long-only, fully invested simplex to minimize:
      sum_i [DEVaRC_i - Dc/N]^2
6. Validate both sum_i(DEVaRC_i) = Dc and DEVaRC_i = Dc/N within tolerance.
7. Record the pure ERC weights separately.
8. For the overlay variant, compare each current contribution with its available positive contribution history, identify excess relative contribution shocks, apply h_i = 1/(1+s_i), multiply the ERC weights by h_i, and renormalize.
9. Rebalance at the same month-end close using MOC execution and observed adjusted closes. Skip orders lacking a finite execution close.
10. Check gross exposure against the 4.0 maximum leverage constraint and proportionally reduce or skip infeasible orders.

Positions are held until the next monthly rebalance. Drift-adjusted one-way turnover is 0.5 × sum_i |w_target,i − w_drifted,i|.