Under Gaussian returns, EVaR-deviation parity reduces to conventional volatility parity. Choi proves the result in Corollary 1 and states it in the abstract. That should be a builder's starting point. Any extra value in the framework must come from the fitted return law's non-Gaussian component. With 25 basis points of cost, matched EVaR versus CVaR Sharpe differences range from -0.292 for XASSET IRP ICA+CTS to +0.021 for SECTOR ICA+NTS ERC. None has p below 0.10.
The construction Choi can actually compute
Risk parity assigns capital by risk contribution rather than expected return. Because EVaR is coherent and positively homogeneous, Euler decomposition gives each asset a share of portfolio EVaR. Choi's contribution is to make those shares computable under tempered stable returns.
In the multivariate normal tempered stable (MNTS) specification, linear aggregation leaves the portfolio NTS. Its EVaR follows from the projected parameters. The asset-level Euler contribution has a closed form comprising a location term, the time-change loading and a diffusion-covariance term scaled by an exponentially tilted factor. No simulation required. Under the independent component route, NTS or CTS components are fitted separately. The contribution then uses the derivative of the component cumulant-generating function at the portfolio tilt.
The distinction between the two tail measures is easy to see under ICA. EVaR allocates through an Esscher-type tilted component mean. CVaR instead uses the component mean conditional on the portfolio's lower-tail event. The mixing matrix stays the same; the component statistic changes.
Choi next removes the fitted location term. Raw EVaR is translation equivariant, so minus the fitted mean appears in its contributions. EVaR deviation removes that term. Under Gaussian returns, the deviation equals sqrt(-2 ln eta) times portfolio volatility. The constant drops out of both inverse-risk normalisation and the equal-contribution condition. Gaussian EVaR-deviation IRP and ERC therefore recover conventional volatility IRP and ERC weights. This corollary makes the paper's best unit test for a replicator: fit a normal, then check that the ERC solver returns the volatility ERC weights. A different answer means the solver is wrong.
The empirical exercise uses monthly, long-only, fully invested portfolios with eta = 0.05. Parameters come from the previous twelve months of daily returns. Choi computes contributions with central finite differences while holding fitted parameters fixed across perturbations. Three universes are studied: seven cross-asset ETFs (XASSET, Apr 2006 to Mar 2026 out of sample), ten French 12-2 momentum deciles (MOM10, from Jan 1996), and nine to eleven SPDR sector ETFs (SECTOR, from Jan 2000). Sharpe differences use a studentized circular block bootstrap with B = 4,999 and block length floor(T^(1/3)). Tests are pairwise and, as the paper states, not adjusted for multiple testing.
The headline comparison is against equal weight. EVaR-ERC has a higher Sharpe in all nine market-specification pairs, while EVaR-IRP does so in eight of nine. Nine of the eighteen differences have p < 0.10. XASSET MNTS produces the largest results: IRP +0.177 (p = 0.016) and ERC +0.246 (p = 0.011). The rest of those rows changes the interpretation. CAGR trails equal weight by 1.50 and 1.52 percentage points, while volatility is lower by 5.68 and 6.37 points. The Sharpe ratio is rewarding de-risking.
Does EVaR win on a fair match?
Choi runs the relevant comparison with the fitted distribution and parity rule held constant, changing only the risk measure. Two of eighteen matched differences reach p < 0.10, both in MOM10 ICA ERC. ICA+NTS records 0.595 versus 0.580 (+0.015, p = 0.028), and ICA+CTS records 0.600 versus 0.582 (+0.018, p = 0.065). The other sixteen matched comparisons have p of 0.10 or above.
The ICA ERC differences also reverse direction across universes. They are +0.028 and +0.017 in SECTOR, compared with -0.061 and -0.126 in XASSET. None of the four has p < 0.10.
One row of Table 3 deserves separate attention. MOM10 ERC is the only place where the Gaussian comparison approaches significance, and the direction varies by model. MNTS loses by 0.007 (p = 0.073). ICA+NTS and ICA+CTS gain +0.014 (p = 0.041) and +0.019 (p = 0.057).
Choi gives the broader conclusion directly: these findings "do not indicate a uniform performance advantage of EVaR over CVaR and should be interpreted within the investment universes, sample periods, and backtest design considered here." The scope matters. These universes contain seven names in XASSET, nine to eleven in SECTOR and ten in MOM10. Cross-sections that small mechanically compress parity differences. The results show it: in SECTOR and MOM10, the Gaussian benchmark comes within 0.01 to 0.03 Sharpe of the tempered stable versions.
Turnover consumes the edge
ICA+NTS EVaR-ERC portfolios trade two to three times as much as their matched CVaR twins. One-way turnover per rebalance is 6.93% against 2.93% in SECTOR, 5.27% against 1.96% in MOM10, and 8.52% against 3.98% in XASSET. The ICA+CTS difference is wider, closer to four times: 9.49% against 2.52% in SECTOR, 8.09% against 1.91% in MOM10, and 12.76% against 3.02% in XASSET.
The cost grid ends at 25 basis points. At that charge, no matched EVaR-CVaR Sharpe difference has p < 0.10. MOM10 ICA+NTS drops from +0.015 (p = 0.028) to +0.010 (p = 0.140). ICA+CTS falls from +0.018 (p = 0.065) to +0.009 (p = 0.373). The Gaussian benchmark comparison follows the same path. The two significant gross gains for MOM10 ERC, +0.014 (p = 0.041) and +0.019 (p = 0.057), contract to +0.008 (p = 0.218) and +0.009 (p = 0.370).
MOM10 also has the least informative cost path because the French decile portfolios are not tradeable. Their internal rebalancing costs remain outside the backtest entirely.
Drawdowns tell a similar story. XASSET ICA+CTS EVaR-IRP posts a 47.11% maximum drawdown, against 23.15% for matched CVaR-IRP. The IRP pair doubles it. For ERC, the figures are 34.56% against 20.93%, roughly 1.65 times. A tail measure producing the larger drawdown delivers the outcome it was meant to prevent.
Choices left to the implementer
Much of the specification is exact: universes, dates, eta, window, rebalance, the finite-difference construction, bootstrap settings and cost grid. Three missing choices determine the weights. We did not find a reported finite-difference step size, an ERC solver tolerance, or a convergence diagnostic. We did not find sensitivity analysis for the twelve-month window or eta. Nor did we find an out-of-sample rule for selecting among MNTS, ICA+NTS and ICA+CTS.
All three are reported, with a different leader in each universe. MNTS ERC wins XASSET at 0.824, ICA+NTS wins SECTOR at 0.588, and ICA+CTS wins MOM10 at 0.600. The paper provides no advance selection rule. Taking the best specification for each universe amounts to choosing in hindsight.
The reader must also supply the MNTS estimation routine. Estimating shape, skew, scale, location and a latent correlation matrix from roughly 250 daily observations is noisy. The contribution formula in equation (41) depends on the fitted diffusion covariance and time-change loading. Estimator choice therefore passes directly into the weights.
Our run lost 3.59% over four and a half years
We implemented the direct-MNTS EVaR-deviation ERC variant using the paper's seven XASSET ETFs. Signals are formed at month-end, with execution at the next session's open. We charged commissions of $0.0040 per share subject to a $1.00 order minimum. Our sample runs from 2020-01-01 to 2024-07-01. These are our figures, not Choi's: cumulative return -3.59%, Sharpe -0.08, maximum drawdown 27.70%, annualised volatility 10.66%, and beta to SPY 0.22.
Choi's corresponding row is XASSET MNTS EVaR-deviation ERC, with 183.78% cumulative return, Sharpe 0.821 and drawdown 20.74% over Apr 2006 to Mar 2026, gross. Our -0.08 versus his 0.821 leaves roughly 0.9 Sharpe between the measurements. Our window is less than a quarter as long and includes commissions. The samples differ most. Our 4.5 years include March 2020 and the 2022 joint equity-duration selloff, without a long benign stretch to average against. For an unlevered long-only cross-asset book, this is the worst available slice.
The risk figures fill in the picture. Beta of 0.22 and volatility of 10.66% fit a rule that loads the low-tail-risk sleeves, AGG and TIP. During 2022, the bond sleeve declined alongside the equity and REIT sleeves. The objective choice contributes little: Choi's own XASSET MNTS comparison gives 0.821 for deviation and 0.824 for raw EVaR. Costs and our open-versus-close fill lower the result, though at roughly 3% monthly turnover neither explains a 0.9-Sharpe gap.
We cannot fully explain the drawdown. Our 27.70% over four and a half years exceeds Choi's 20.74% over twenty, suggesting that our realised weights were more concentrated. Most periods were positive, 63.48% of them, and gross gains covered gross losses 1.15 to 1. The overall loss came from a small number of large down months.
One candidate is the solver. Our 1e-5 central-difference step and our own convergence criterion are implementation choices, while the paper supplies no benchmark Euler error against which to check them. Our result is evidence from our single run on a hostile window, and not a verdict on the paper.
Keep the contribution formulas
The derivations will last. Choi gives asset-level Euler contributions for EVaR under both representations, along with a location separation that lets the user decide whether the fitted mean belongs in a risk budget. A finite-difference fallback handles the weight-dependent admissible-domain endpoint, although the paper notes that it fails at nonsmooth switching points. The Gaussian corollary also deserves to be written down.
We previously examined tempered-stable calibration in a very different setting, a diffusion denoiser that ultimately did not use it (note). Here the calibration feeds the weights, leaving the estimation questions active.
I would move a production risk budget from CVaR to EVaR on one condition: a matched comparison on a universe large enough to avoid mechanically tiny parity differences, using a pre-committed model specification and the same turnover as the CVaR book. The advantage would still need significance at 25 basis points. Across seven to eleven names, none of Choi's three universes meets that standard.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.