AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Long-only worst-case tracking edges 1/N without an index comparison

Across 25 large caps in 2020-2024, uncertain moments cut the leading monthly mean from 3.38% to 2.33%.

7 min read · Enhanced index tracking / robust portfolio optimization · US equities

Reviewing: Robust enhanced index tracking portfolio selection under distributional uncertainty · Jun Cai and Zhiqiao Song · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Long-Only Robust Enhanced Index Tracking with Daily-Moment M-ER

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (7)

  • Equation 5 monthly adjusted-close stock-loss definition and next-out-of-sample-month weighted return.: Estimate daily adjusted-close losses and begin applying monthly decisions at the next trading day's observed open; calculate realized returns from actual execution prices rather than treating decision-close returns as executable. (invalidates: Direct numerical comparability with the paper's Table 1 and Table 2; direct numerical comparability with Figure 1(d) and Figures 5–7)
  • Paper universe: the 25 named stocks, five from each of five sectors.: Select 25 stocks anew for each calendar year by year-specific aiquant_screening_table capitalization. (invalidates: Direct numerical comparability with the paper's Table 1 and Table 2; direct numerical comparability with Figure 1(d) and Figures 5–7)
  • Paper monthly losses, rolling 60-month estimation, monthly refitting, January 2015–December 2019 initial training and January 2020–December 2024 evaluation.: Retain a trailing 60-calendar-month window and monthly refitting but estimate its moments from daily losses; evaluate January 2020–July 1, 2024. (invalidates: Direct numerical comparability with the paper's Table 1 and Table 2; direct numerical comparability with Figure 1(d) and Figures 5–7)
  • Paper α=−0.04 per month, r=−0.01 per month, γ1=0.22, γ2=0.025 and κ=5 for its generalized-uncertainty experiment.: Keep the quoted monthly values as references but optimize daily losses using α_daily=−0.04/21, r_daily=−0.01/21, γ1_daily=0.22/21 and γ2_daily=0.025/21; retain κ=5. (invalidates: Direct numerical comparability with the paper's Table 2; direct numerical comparability with Figures 5–7)

3 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 62.4%Jul 2024
Sharpe
0.70
Total Return
62.4%
Max Drawdown
-32.5%
CAGR
11.6%
Volatility
20.1%
Beta vs SPY
0.90
Trades
1,304

What the paper reports for its own strategy

  • M-ER, known mean/covariance, short selling allowed, OOS Jan 2020-Dec 2024, monthly rebalancing, transaction costs not stated: mean return 0.038021, SD 0.110156, Sharpe 0.345155, Sortino 0.493813, Omega 0.954801 (Sharpe not stated as annualized; risk-free rate not stated)
  • M-TSV, known mean/covariance, short selling allowed, same period, costs not stated: mean return 0.048939, SD 0.150532, Sharpe 0.325107, Sortino 0.500902, Omega 1.169271
  • M-ER, known mean/covariance, no short selling, same period, costs not stated: mean return 0.027183, SD 0.064414, Sharpe 0.422011, Sortino 0.508637, Omega 0.608300
  • M-TSV, known mean/covariance, no short selling, same period, costs not stated: mean return 0.033821, SD 0.074642, Sharpe 0.453109, Sortino 0.611352, Omega 0.817288
  • M-ER, generalized set D(0.22, 0.025), short selling allowed, same period, costs not stated: mean return 0.023264, SD 0.053799, Sharpe 0.432424, Sortino 0.480342, Omega 0.456294
  • M-ER, generalized set, no short selling, same period, costs not stated: mean return 0.023280, SD 0.052341, Sharpe 0.444768, Sortino 0.489637, Omega 0.447392

Automated analysis of the paper

A 3.38% monthly long-only winner falls to 2.33% when its estimated moments can be wrong. Equal weighting earns 2.03% in that second test. Cai and Song never put either result against the index their portfolios are meant to beat.

The index enters through the penalty

The paper measures losses, where a positive number is bad. Its active loss is portfolio loss minus index loss, then minus α; −α is the desired excess return. The authors choose 4% a month (−α = 0.04). The in-sample average excess return was 0.75%, leaving that target more than five times the observed average.

The optimizer combines worst-case mean portfolio loss with κ times a worst-case penalty for falling short of index-plus-4%. At m = 1, the penalty averages the shortfall, or expected regret. At m = 2, it averages squared shortfall, or target semi-variance. The combined models are M-ER and M-TSV; ER and TSV drop the mean-loss term. All carry a floor on worst-case expected return. The Pareto-optimal result follows directly: a minimizer of a weighted sum of two objectives cannot be improved on both objectives at once.

The first uncertainty set treats the stock and index mean vector and covariance as known. Across every distribution with those moments, worst-case active loss depends only on its mean and standard deviation. TSV has an exact closed form: the squared positive part of mean shortfall plus tracking variance. Its downside penalty therefore trades mean shortfall against tracking error. For ER with short selling, a solution exists only above κ0, a threshold of 4.65 in the in-sample data. The authors solve the long-only versions as convex programs and use κ = 5 in the main tests.

The second set comes from Kang et al. (2019). It allows means to move within an ellipsoid of size γ1 = 0.22 and covariance within a Frobenius ball of size γ2 = 0.025, while holding the index mean fixed. Worst-case mean loss becomes sample mean plus √γ1 times portfolio volatility. Roughly 0.47 of portfolio SD is thus added to the mean loss, favouring lower-volatility weights. These versions, too, are solved as convex programs.

High trailing 60-month sample means drive the return side of the objective. The tracking penalty, scaled by κ, reins those positions in.

One universe across five years

The authors take the five largest companies by market capitalization from each of five sectors: technology, healthcare, financials, consumer cyclical and consumer defensive. The resulting 25 names include NVDA, AVGO, LLY, TSLA and BKNG. Monthly losses use Yahoo! Finance adjusted closes from January 2015 to December 2024. Each fit uses a rolling 60-month window. Monthly out-of-sample returns run from January 2020 to December 2024, giving 60 returns in tests with and without short selling.

The Sharpe column appears to divide monthly mean by monthly SD without subtracting a risk-free rate; the paper does not specify the convention. For M-TSV, 0.033821 divided by 0.074642 gives the 0.453 reported in Table 1. The same fixed list is applied to every window from 2015 to 2024, with no selection date supplied. A ranking taken near 2024 would give the 2020-2024 test survivorship.

Where is the index?

Each results table has seven rows: four tracking models, worst-case mean-variance (M-V), worst-case variance (V) and equal weight (EWP). Neither includes the index. We could not find an identification of index X0 in the empirical section, or a reported tracking error or excess return against it. ER and TSV appear in the tables and, by the paper's definition, measure shortfall against index-plus-target. Those are the sole index-relative figures. Their underlying index remains unnamed.

The abstract claims outperformance of "the tracked index". The cumulative-wealth figures might have settled the comparison, yet their captions list models (a)-(g) and show no index line. The conclusion instead repeats the outperformance claim against EWP, robust M-V and robust V. For a portfolio targeting 4% a month above its benchmark, an index-relative shortfall measure with no named index leaves the central comparison unresolved.

The long-only lead contracts

Under known moments and without short selling, M-TSV records a 0.033821 monthly mean, 0.074642 SD and 0.453 Sharpe in Table 1. EWP records 0.020299, 0.050758 and 0.400. Allow uncertain moments in the long-only test and M-ER leads Table 2 instead, at 0.023280, 0.052341 and 0.445. All seven long-only means in that table range from 0.016281 for V to 0.023280 for M-ER, with EWP inside the range. The monthly return advantage over EWP contracts from 1.35 points to 0.30; the Sharpe advantage moves from 0.053 to 0.045. A monthly Sharpe based on 60 observations has a standard error near 0.13 (1/√60). We found no significance test in the paper.

Short selling makes the known-moment Sharpe comparison less flattering. M-TSV has 0.150532 monthly SD and a 0.325 Sharpe in that Table 1 block. TSV reaches 0.397, while EWP leads at 0.400. The authors acknowledge that, with shorts, "the Sharpe and Sortino ratios are improved with models that minimize downside risk alone." The Sortino statement holds for the ER pair, 0.505900 versus M-ER's 0.493813. For the other pair, M-TSV's 0.500902 exceeds TSV's 0.457940.

The abstract nevertheless says the combined models "consistently outperform" the downside-only models. The authors rely on mean return, Omega and cumulative wealth for that claim. M-TSV's 0.048939 mean beats TSV's 0.023855, alongside monthly SD three times EWP's 0.050758. The text calls M-TSV the Omega leader in this with-shorts block, although Table 1 reports 1.189569 for M-V and 1.169271 for M-TSV.

The κ recommendations deserve separate caution. Setting the headline κ = 5 from an in-sample threshold is fair. The reported "optimal" κ* values, however, come from scanning out-of-sample outcomes. With short selling and known moments, Sortino peaks at 7.1, 7.5 and 8.1. For long-only portfolios under known moments, Sharpe peaks at 2.1 for M-TSV and 1.7 for M-V. The authors concede κ "does not have a straightforward empirical approach for its determination." The conclusion still presents those values as "empirical guidance for selecting optimal weighting coefficients to maximize Sharpe, Sortino, and Omega ratios." They describe what won in the test window.

Our run uses a different portfolio

We ran a long-only, uncertain-moment M-ER backtest at κ = 5 from January 2020 to July 2024. Several choices differ from the paper:

Our run returned 62.44%, for an 11.6% CAGR. Sharpe was 0.70, Sortino 0.87, volatility 20.11% and maximum drawdown −32.50%. The closest paper result is long-only M-ER on the generalized set, without short selling in Table 2: a 0.0233 monthly mean and 0.445 monthly Sharpe. Those figures annualize to roughly 28% a year and a Sharpe near 1.5. Ours is lower by more than half on both measures, across setups that hold different portfolios. The 28% multiplies the monthly mean by twelve, whereas our 11.6% is compounded; variance drag therefore makes the arithmetic comparison somewhat closer. The paper's 0.0523 monthly SD annualizes to about 18%, near our 20.11%. Most of the gap is in return.

Our 10% cap takes weight away from the high-mean stocks the objective prefers. Re-screening the universe at the time probably removes some of the paper's largest early winners. Dividing the mean by 21 while volatility falls by √21 may also make κ = 5 more conservative. Our window finishes six months earlier, and commissions impose a small drag. If our Sharpe subtracted a risk-free rate near 5% in 2023-24, its definition could explain part of the difference as well. Together, these possibilities account for only part of the gap. Our beta to the S&P 500 was 0.90. We did not calculate excess return or tracking error against it, so our run leaves the index question open too. It is one automated pass under different assumptions, rather than a verdict on the authors' results.

The comparison I would need

The paper's fixed rule-based 25 selects the top five by capitalization in five chosen sectors, without a stated selection date. We found no turnover or trading-cost figure for monthly rebalancing. Known-moment combined portfolios with shorts also run at 0.110 to 0.155 monthly SD. A table showing generalized-set M-ER against a named index over 2020-2024, with excess return and tracking error, would change my view.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each calendar-year entry, select that year's 25 largest eligible non-ADR STOCK rows.
At each month-end close, align observed stock and SPX closes over the trailing 60 calendar months; compute daily losses and their joint sample mean and covariance.
If benchmark coverage or synchronized observations are insufficient, skip the decision.
Solve generalized-moment-uncertainty M-ER (equation 4.7, m=1) with kappa=5, the calibrated daily thresholds, nonnegative weights summing to one, a 10% stock cap, and the worst-case return constraint. Include the index loss in benchmark-relative risk; assign the index no holding weight.
If infeasible, skip the M-ER rebalance rather than substitute weights. Record equal-weight, robust M-V, and gamma1-sensitivity comparators separately.
Trade the selected weights at the run's market-on-close auction; account for the platform's per-fill commissions.