Automated analysis of the paper
A 3.38% monthly long-only winner falls to 2.33% when its estimated moments can be wrong. Equal weighting earns 2.03% in that second test. Cai and Song never put either result against the index their portfolios are meant to beat.
The index enters through the penalty
The paper measures losses, where a positive number is bad. Its active loss is portfolio loss minus index loss, then minus α; −α is the desired excess return. The authors choose 4% a month (−α = 0.04). The in-sample average excess return was 0.75%, leaving that target more than five times the observed average.
The optimizer combines worst-case mean portfolio loss with κ times a worst-case penalty for falling short of index-plus-4%. At m = 1, the penalty averages the shortfall, or expected regret. At m = 2, it averages squared shortfall, or target semi-variance. The combined models are M-ER and M-TSV; ER and TSV drop the mean-loss term. All carry a floor on worst-case expected return. The Pareto-optimal result follows directly: a minimizer of a weighted sum of two objectives cannot be improved on both objectives at once.
The first uncertainty set treats the stock and index mean vector and covariance as known. Across every distribution with those moments, worst-case active loss depends only on its mean and standard deviation. TSV has an exact closed form: the squared positive part of mean shortfall plus tracking variance. Its downside penalty therefore trades mean shortfall against tracking error. For ER with short selling, a solution exists only above κ0, a threshold of 4.65 in the in-sample data. The authors solve the long-only versions as convex programs and use κ = 5 in the main tests.
The second set comes from Kang et al. (2019). It allows means to move within an ellipsoid of size γ1 = 0.22 and covariance within a Frobenius ball of size γ2 = 0.025, while holding the index mean fixed. Worst-case mean loss becomes sample mean plus √γ1 times portfolio volatility. Roughly 0.47 of portfolio SD is thus added to the mean loss, favouring lower-volatility weights. These versions, too, are solved as convex programs.
High trailing 60-month sample means drive the return side of the objective. The tracking penalty, scaled by κ, reins those positions in.
One universe across five years
The authors take the five largest companies by market capitalization from each of five sectors: technology, healthcare, financials, consumer cyclical and consumer defensive. The resulting 25 names include NVDA, AVGO, LLY, TSLA and BKNG. Monthly losses use Yahoo! Finance adjusted closes from January 2015 to December 2024. Each fit uses a rolling 60-month window. Monthly out-of-sample returns run from January 2020 to December 2024, giving 60 returns in tests with and without short selling.
The Sharpe column appears to divide monthly mean by monthly SD without subtracting a risk-free rate; the paper does not specify the convention. For M-TSV, 0.033821 divided by 0.074642 gives the 0.453 reported in Table 1. The same fixed list is applied to every window from 2015 to 2024, with no selection date supplied. A ranking taken near 2024 would give the 2020-2024 test survivorship.
Where is the index?
Each results table has seven rows: four tracking models, worst-case mean-variance (M-V), worst-case variance (V) and equal weight (EWP). Neither includes the index. We could not find an identification of index X0 in the empirical section, or a reported tracking error or excess return against it. ER and TSV appear in the tables and, by the paper's definition, measure shortfall against index-plus-target. Those are the sole index-relative figures. Their underlying index remains unnamed.
The abstract claims outperformance of "the tracked index". The cumulative-wealth figures might have settled the comparison, yet their captions list models (a)-(g) and show no index line. The conclusion instead repeats the outperformance claim against EWP, robust M-V and robust V. For a portfolio targeting 4% a month above its benchmark, an index-relative shortfall measure with no named index leaves the central comparison unresolved.
The long-only lead contracts
Under known moments and without short selling, M-TSV records a 0.033821 monthly mean, 0.074642 SD and 0.453 Sharpe in Table 1. EWP records 0.020299, 0.050758 and 0.400. Allow uncertain moments in the long-only test and M-ER leads Table 2 instead, at 0.023280, 0.052341 and 0.445. All seven long-only means in that table range from 0.016281 for V to 0.023280 for M-ER, with EWP inside the range. The monthly return advantage over EWP contracts from 1.35 points to 0.30; the Sharpe advantage moves from 0.053 to 0.045. A monthly Sharpe based on 60 observations has a standard error near 0.13 (1/√60). We found no significance test in the paper.
Short selling makes the known-moment Sharpe comparison less flattering. M-TSV has 0.150532 monthly SD and a 0.325 Sharpe in that Table 1 block. TSV reaches 0.397, while EWP leads at 0.400. The authors acknowledge that, with shorts, "the Sharpe and Sortino ratios are improved with models that minimize downside risk alone." The Sortino statement holds for the ER pair, 0.505900 versus M-ER's 0.493813. For the other pair, M-TSV's 0.500902 exceeds TSV's 0.457940.
The abstract nevertheless says the combined models "consistently outperform" the downside-only models. The authors rely on mean return, Omega and cumulative wealth for that claim. M-TSV's 0.048939 mean beats TSV's 0.023855, alongside monthly SD three times EWP's 0.050758. The text calls M-TSV the Omega leader in this with-shorts block, although Table 1 reports 1.189569 for M-V and 1.169271 for M-TSV.
The κ recommendations deserve separate caution. Setting the headline κ = 5 from an in-sample threshold is fair. The reported "optimal" κ* values, however, come from scanning out-of-sample outcomes. With short selling and known moments, Sortino peaks at 7.1, 7.5 and 8.1. For long-only portfolios under known moments, Sharpe peaks at 2.1 for M-TSV and 1.7 for M-V. The authors concede κ "does not have a straightforward empirical approach for its determination." The conclusion still presents those values as "empirical guidance for selecting optimal weighting coefficients to maximize Sharpe, Sortino, and Omega ratios." They describe what won in the test window.
Our run uses a different portfolio
We ran a long-only, uncertain-moment M-ER backtest at κ = 5 from January 2020 to July 2024. Several choices differ from the paper:
- We re-screened a top-25 non-ADR universe each year using that year's capitalization.
- We used the S&P 500 for X0.
- We used daily losses over a trailing 60 months and divided the paper's monthly parameters by 21.
- We capped each stock at 10%.
- We filled month-end decisions at the closing auction and charged $0.004 a share in commissions.
Our run returned 62.44%, for an 11.6% CAGR. Sharpe was 0.70, Sortino 0.87, volatility 20.11% and maximum drawdown −32.50%. The closest paper result is long-only M-ER on the generalized set, without short selling in Table 2: a 0.0233 monthly mean and 0.445 monthly Sharpe. Those figures annualize to roughly 28% a year and a Sharpe near 1.5. Ours is lower by more than half on both measures, across setups that hold different portfolios. The 28% multiplies the monthly mean by twelve, whereas our 11.6% is compounded; variance drag therefore makes the arithmetic comparison somewhat closer. The paper's 0.0523 monthly SD annualizes to about 18%, near our 20.11%. Most of the gap is in return.
Our 10% cap takes weight away from the high-mean stocks the objective prefers. Re-screening the universe at the time probably removes some of the paper's largest early winners. Dividing the mean by 21 while volatility falls by √21 may also make κ = 5 more conservative. Our window finishes six months earlier, and commissions impose a small drag. If our Sharpe subtracted a risk-free rate near 5% in 2023-24, its definition could explain part of the difference as well. Together, these possibilities account for only part of the gap. Our beta to the S&P 500 was 0.90. We did not calculate excess return or tracking error against it, so our run leaves the index question open too. It is one automated pass under different assumptions, rather than a verdict on the authors' results.
The comparison I would need
The paper's fixed rule-based 25 selects the top five by capitalization in five chosen sectors, without a stated selection date. We found no turnover or trading-cost figure for monthly rebalancing. Known-moment combined portfolios with shorts also run at 0.110 to 0.155 monthly SD. A table showing generalized-set M-ER against a named index over 2020-2024, with excess return and tracking error, would change my view.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.