Cai, Mao and Song's wealth leader takes more tail risk. With mean and covariance treated as known, their reward-penalty CVaR rule selects a point on the Markowitz frontier. Its wealth lead over the comparators largely reflects which point it selects. The theory earns attention; the empirical claim has to be judged against that frontier.
The bonus clause changes the loss
The fund manager earns a reward at rate β when the portfolio return exceeds −d. A loss beyond d triggers a penalty at rate λ, which the authors liken to a clawback. For the institution, the resulting loss is (1−β)X + (β−λ)(X−d)+ plus a constant. With λ < β, dividing by (1−β) gives X + κ(X−d)+, where κ = (β−λ)/(1−β). It is portfolio loss with a weighted stop-loss on the same portfolio. The authors also let κ stand as a free weight on downside risk, with d set at a risk level such as a VaR or CVaR.
Theorem 3.1 computes the worst-case CVaR of this payoff across distributions with mean µ and standard deviation σ. A three-point distribution attains the worst case. For λ < β, the result has three regimes, determined by d−µ relative to σd1 and σd2. In the middle regime, it is (1−β)d + [Δ − (2−λ−β)(d−µ)]/(2(1−α)), with Δ built from σ² + (d−µ)². The formula contains earlier results as special cases:
- Chen et al. (2011)'s µ + σ√(α/(1−α)), at λ = β = 0.
- Cai et al. (2024)'s worst-case stop-loss CVaR, at λ = 0 and β = 1, here extended to any real d.
- Jagannathan (1977)'s worst-case stop-loss premium, as α → 0, with λ = 0 and β = 1.
The paper also finds a distinction absent from the usual mean-variance set: worst-case VaR and worst-case CVaR, known to coincide there, separate when λ < β.
Which frontier point gets picked?
A univariate projection lemma reduces each portfolio's worst case to the scalar problem. Theorem 4.1 puts the optimal weights on the mean-variance frontier at a mean loss s. Choosing the portfolio comes down to choosing s: it minimizes a strictly convex g(s) and is unique under the technical condition (4.6). The budget constraint is ω⊤e = 1, with weights otherwise free in Rⁿ. Shorting and leverage are allowed. We found no position limits.
The authors call the α → 0 version M-WC-EDR: mean loss plus κ times the worst-case expected stop-loss. Its s has the explicit expression (4.8), provided κ > κ = (2+2√(b0+1))/b0. At κ ≤ κ, including κ = 0, there is no solution. The denominator in (4.8) contains 4(κ+1) − b0κ², a factor that reaches zero exactly at κ. As κ approaches κ from above, s falls without bound. Expected return and variance for the chosen portfolio both rise without limit. For M-WC-EDR, the paper's advice to wealth-maximizing, risk-seeking investors to choose κ just above κ* therefore means taking ever more leverage on the frontier.
Kang et al. (2019)'s generalized set allows the mean to move within an ellipsoid of size γ1 and the covariance to move within Frobenius distance γ2. The worst case raises portfolio mean loss by √γ1 times the portfolio standard deviation under Σ̂; portfolio variance becomes ω⊤(Σ̂ + γ2·I)ω. This yields the convex program (5.2). Explicit solutions are generally unavailable, and the Figure 3 portfolios do not lie on the frontier of (µ̂, Σ̂).
Wealth and tail risk move together
The study follows 15 US stocks, the three largest in each of five sectors, using daily Yahoo data from March 2017 to February 2025. Training takes the first three years; the test runs from March 2020 to February 2025. A 756-day window supplies the moments and rolls forward one day at a time. The settings are α = 0.9, a mean-variance target of r0 = 0.0003 daily and κ = 0.5. At the three targets −d = 0.006, 0.0003 and 0.00006, M-WC-EDR's cumulative wealth is "significantly higher" than every comparator's for most of the test period. That is the paper's description of a plot, without a test behind it. M-WC-EDR also has the highest worst-case CVaR. WC-CVaR-BPL, the paper's κ-weighted CVaR model, earns a slight but consistent cumulative-wealth lead over plain worst-case CVaR.
The authors acknowledge a "trade-off between improving expected portfolio return and controlling the worst-case portfolio CVaR" in the abstract. They also say that adding downside risk to portfolio loss "can achieve higher investment returns than considering either downside risk or portfolio loss alone." The same sentence claims better risk management. Yet Figure 1 gives M-WC-EDR, the wealth leader, the highest worst-case CVaR of the six. Theorems 4.1 and 4.2, together with the same mean-variance reduction for the comparators, make that risk-management claim difficult to treat as an edge. Every Figure 1 model chooses a point on the same frontier. Higher return accompanied by higher worst-case CVaR places a portfolio further along it. A desk would want the models compared at matched worst-case CVaR, or by Sharpe ratio. We found neither. The paper prints no return, Sharpe or drawdown figures; every empirical result appears as a plot.
The plotted CVaR is model-implied sup CVaR, computed using each window's sample moments. The paper calls its out-of-sample series "expected portfolio returns" constructed from "the expected daily loss vector". We could not determine from the text whether the wealth series compounds realized daily returns.
The mean-variance benchmark makes the choice of frontier point especially visible. It produces the lowest wealth. Its worst-case CVaR falls below M-WC-EDR's, though it sits slightly above the other four models'. An equality return constraint fixes its position at the authors' chosen r0 = 0.0003 target.
The κ sweep tells a similar story. At 1, 5, 10, 15, with d = 0.03, smaller κ brings M-WC-EDR and M-WC-CVaR more wealth and more CVaR. WC-CVaR-BPL moves in the opposite direction.
Test-period information in γ1 and γ2
To set the ambiguity sizes, the authors calculate deviations window by window using "a rolling window approach over all the data". All the data includes 2020 to 2025, so test-period information enters γ1 = 0.009 and γ2 = 0.001. Covariance ambiguity is particularly consequential here: D_n(0, 0.001) produces the highest wealth, and γ2 matters more than γ1. Figure 3 lists (0.009, 0.001) twice among its settings, while the text makes a comparison with D_n(0.009, 0). Its daily return floor, r = −0.098, looks unlikely to bind.
We found no date for the market-cap ranking. The inclusion of NVDA and LLY makes the list read like a recent one; if it is recent, the universe incorporates test-period knowledge. Nor did we find a mention of transaction costs, despite daily rebalancing. The authors are candid about κ: "it is unlikely that a single value of κ can simultaneously satisfy both objectives". They leave a selection criterion to future work. We also found no reported b0, leaving readers unable to check whether κ = 0.5 exceeds κ* for models that require it.
Our long-only run, 2020 to mid-2024
We built one version from the paper's description, in one automated pass. It traded daily from 2020-01-01 to 2024-07-01 across 14 of the paper's stocks, excluding GOOG. We used the classical fixed-moment objective, α = 0.90, β = 0.5, λ = 0 (so κ = 1) and d = 0.03, with a 756-day window. We required long-only weights, capped each name at 10%, and charged $0.004 a share in commissions.
The run returned 77.56% in total, with a 0.97 Sharpe, 17.82% volatility and a −27.78% maximum drawdown over 7,456 trades. The paper supplies no return or Sharpe figure to place beside our 77.56% total return or 0.97 Sharpe. Its setup also differs substantially from ours: the authors allow unconstrained weights, while ours were long-only. With 14 names and a 10% cap, at least 10 names have to carry weight. Our optimizer consequently had little latitude to depart from a near-equal-weight mega-cap book. Its 0.74 beta to SPY suggests it still favored the defensive names within that cap.
There is a fill discrepancy in our run. We intended to trade at the next day's open; the recorded fills occurred at the close. We also did not produce worst-case CVaR (WC-CVaR) or mean-variance (M-V) comparators. This one automated pass is no verdict on the authors' work, and its result cannot establish whether the threshold penalty beats either comparator.
The theory is a real generalization. The empirical case would persuade me if M-WC-EDR beat the other Figure 1 models at matched worst-case CVaR, and if Figure 3 held up with γ1 and γ2 estimated solely on the 2017 to 2020 training window.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.