AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Five surviving tail laws, half a basis point of decision

A truncated NPMLE grid sets the worst case, and the loss-averse investor barely participates

2026-09-08 · 8 min read

Reviewing: Mixing-Law Uncertainty in Multivariate Normal Mean-Variance Mixtures: Semi-parametric Estimation and Robust Cumulative-Prospect Decisions · Nuerxiati Abudurexiti · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Distributionally Robust NMVM-CPT Monthly Equity Exposure on Paper 30 US Stocks

Backtest period 2020-01-01 to 2025-10-08 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (6)

  • : (invalidates: Paper Table 10 c* and gross risky exposure values; paper Figure 4 lower-envelope shape; paper projection diagnostics in Table 12; direct numerical comparability of c_max=0.045458)
  • : (invalidates: Results no longer represent a dynamic top-30-by-prior-year-market-cap strategy; universe turnover from aiquant_screening_table is not tested)
  • Position and leverage risk constraints (max_position_size_pct, max_leverage): max_leverage reduced from 4.0 to 1.0 to match the paper's gross risky-asset bound L=1; a platform per-name cap max_position_size_pct=0.1 is retained although the paper imposes no per-asset cap. (invalidates: None of the paper's reported results — with the restored eq-109 direction and 0<=c<=c_max the gross risky exposure never exceeds L=1, so neither platform cap binds.)
  • Distillation loss of a defining construct: the paper's common direction is the mean-variance fund direction q0 = S_tr^{-1} v_tr / (v_tr^T S_tr^{-1} v_tr) (eq 62/109), and every Table 10 c*/gross-exposure, Table 12 projection diagnostic and Figure 4 lower-envelope shape is defined relative to it; the method block replaced it with 'equal-weight long-only or inverse-volatility long-only normalized to gross exposure one' (signal_construction step 5), which detaches all those numbers — the specification correctly restored eq-109 in portfolio_direction and declared the correction, so the loss is located at the distillation step and is not live in the spec.

2 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 0.1%Oct 2025
Sharpe
0.13
Total Return
0.1%
Max Drawdown
-0.2%
CAGR
0.0%
Volatility
0.1%
Trades
1,230

What the paper reports for its own strategy

  • Model-based lower-envelope CPT value at the robust optimum: -0.04478 at 5% annual reference return; -0.09245 at 10% (training-estimated parameters, 1,024 scenarios)
  • Empirical holdout CPT value at the robust weights: -0.04867 at 5% reference; -0.10049 at 10% reference (holdout 483 days, 30 Jun-type chronological split, no transaction costs) - the authors state these are descriptive, not independent out-of-sample, because the same holdout defined the ambiguity set
  • Prospect values remain negative at both positive reference returns, with robust gross risky exposure of only 4.286% (5%) and 9.769% (10%) of the L=1 gross-exposure budget

With 1,610 daily observations on 30 US large caps, a skewed-t and a lognormal scale mixture remain indistinguishable. A Gaussian does not. That asymmetry is the result from Abudurexiti's paper that matters for a trading desk.

The abstract acknowledges the identification problem directly: "Several parametric and semi-parametric models, however, remain in the ambiguity set." Its next line supplies the defence. The portfolio optimization stage determines the worst-case model, rather than accepting a point estimate from the holdout score in advance. Yet the stage-two decision barely changes. At the 5% annual reference return, the worst-case choice differs from the point-estimate winner, inverse gamma, by 0.051 percentage points of gross weight, 4.286% against 4.337%. At the 10% reference, the difference is 0.080 pp, with 9.769% against 9.849%. In this application, all that machinery has almost no effect.

The mixing law carries the tails

The paper writes the return vector as X = mu + gamma*Z + sqrt(Z)*A*N. The positive latent scale Z is independent of a standard normal vector. It scales the conditional covariance and shifts the conditional mean through gamma. Given Z = z, X is Gaussian, with mean mu + z*gamma and covariance z*Sigma. The law of Z therefore contains the tails and the skew loading. Inverse gamma produces the skewed t, GIG produces generalized hyperbolic, gamma produces variance gamma, and inverse Gaussian produces NIG.

Scale itself cannot identify Z. Dividing Z by s while multiplying gamma and Sigma back leaves the model unchanged. Abudurexiti resolves this by fixing the determinant, |Sigma| = |S_n| from the training covariance, rather than imposing E(Z) = 1. The estimation then includes m = E(Z). Fitted means are 1.007 for inverse gamma, 1.025 for the NPMLE and 0.965 for lognormal. Those estimates of m give the same portfolio different expected excess returns under the fitted laws.

Six parametric laws are fitted by EM/ECM: GIG, inverse Gaussian, inverse gamma, gamma, exponential and lognormal. A grid nonparametric maximum likelihood estimator is fitted alongside them, using 45 log-spaced support points on [0.025, 16]. The sample contains daily log returns in percentage units for 30 large-cap US names from 2 January 2020 to 1 June 2026. There are 1,611 common trading days and 1,610 returns. The chronological split leaves 1,127 observations for training and 483 for holdout.

Mean holdout log score ranks the models. Each model is then compared with the winner through a paired circular block bootstrap, using block length five and 4,000 replications. A specification remains in the finite ambiguity set whenever its 95% interval covers zero.

The bootstrap leaves five standing

Five of eight specifications survive. Inverse gamma leads with a holdout score of -52.53145 and serves as the reference. GIG trails by 0.00002, with interval [-0.00016, 0.00018]. The deficits are 0.00708 for inverse Gaussian, 0.01015 for lognormal and 0.00835 for the NPMLE, whose interval is [-0.00398, 0.02023]. Gamma is rejected at 0.08996, [0.02959, 0.15821], as is exponential at 0.15851, [0.10904, 0.20505]. The multivariate Gaussian falls behind inverse gamma by 4.30438, with interval [3.07520, 5.63994].

GIG's third parameter contributes nothing visible here. Its fitted psi is about zero, the inverse-gamma boundary, while the holdout score gap is 1.9e-5. With one fewer parameter, inverse gamma also wins on AIC, 109483.9 to 109485.9. The broader family adds a penalty without improving the fit.

Equal density scores still leave room for different return laws. In the 30-dimensional fits, inverse gamma and the NPMLE have minimum eigenvalues of 0.428 and 0.429, maximum eigenvalues of 44.206 and 44.124, and mean correlation of 0.362 under both. Their largest absolute off-diagonal correlation difference is 0.0049; the RMS difference is 0.0015. Latent-scale variances are similarly close at 1.7748 against 1.6837. Their material distinction lies in support: inverse gamma is unbounded, while the fitted grid ends at 16.6682.

The paper's claim that the mixing distributions differ much more than the fitted central return densities refers to the AAPL latent-scale estimates, rather than the 30-asset fit. The desk-level reading is simpler. Across these two estimators, the mixing law determines the tail scenarios while the correlation matrix barely changes, with a largest off-diagonal difference of 0.0049.

The marginal analysis of four names, AAPL, AMZN, GOOGL and MSFT, reaches the same issue from another direction. GIG records the best training score at -2.0848. Lognormal leads on holdout at -1.9452 and has the smallest AIC, 18833.0, and BIC, 18913.5. Inverse gamma delivers the smallest average absolute quantile error at the 5% and 95% levels, although its error exceeds the Gaussian's in the 1% left tail. Three of four inverse-gamma alpha estimates land exactly on the imposed lower bound of 2.05, which the author identifies as constrained estimates. Second moments exist there by imposition.

One ray across five laws

The decision stage takes one direction from the training moments:

q0 = S_tr^-1 v_tr / (v_tr' S_tr^-1 v_tr).

Only portfolios of the form x(c) = c*q0 are considered. A gross bound of L = 1 implies c_max = 0.045458. Each surviving law is projected onto this ray to produce a scalar excess return, then assigned a cumulative prospect value function of c. The Tversky-Kahneman parameters are alpha+ = alpha- = 0.88, lambda = 2.25, delta+ = 0.61 and delta- = 0.69. The optimizer chooses the exposure that maximizes the pointwise minimum of the five curves. Scenario generation uses 2^17 scrambled Sobol draws, compressed into 1,024 equally weighted quantile points.

At a 0% annual reference return, every model selects c = 0. With the reference at 5%, the worst-case exposure is c = 0.001948, corresponding to 4.286% gross risky weight. At 10%, c* rises to 0.004441, or 9.769%. Table 10 reports model-specific exposures from 4.272% to 4.349% at the 5% reference and from 9.738% to 9.912% at 10%. The NPMLE alone is active at both positive reference gaps. Its answer and the worst-case answer therefore coincide to six decimals.

The author's conclusion describes the model-specific exposures as close. Table 10 confirms it. Among the five laws retained by the formal bootstrap comparison, the full gross-weight range is 0.077 percentage points at the 5% annual reference and 0.174 percentage points at the 10% reference.

The pessimistic law is also the one with a fitted scale distribution that ends. The grid NPMLE uses 24 effective support points, with the largest at 16.6682 and no mass beyond it. Inverse gamma has unbounded support. The paper reports that the NPMLE is uniquely active at both positive reference gaps and, separately, that the grid NPMLE depends on its specified support range. We read the truncated grid as the source of the NPMLE's pessimism here. Anyone adopting the method should test the grid endpoints first.

A stronger economic result lies beneath the ambiguity exercise. Prospect value at the optimum equals -0.04478 for a 5% reference and -0.09245 for 10%. With loss aversion and probability weighting, the investor uses roughly 4% to 10% of an available gross budget of one in a 30-stock large-cap book, yet values the resulting position as a loss. The CPT preference parameters remain fixed at published experimental estimates, alpha 0.88, lambda 2.25, delta+ 0.61 and delta- 0.69. Their sensitivity is left untested, while the reference return varies continuously from 0% to 15%.

Our monthly walk-forward run

We ran a monthly walk-forward version on the paper's 30 tickers from 1 January 2020 to 8 October 2025. The result is flat: Total Return 0.10%, Sharpe 0.13, Sortino 0.15, Calmar 0.07, Max Drawdown -0.25% and Profit Factor 1.07. Over roughly 5.8 years, a 0.10% total return with a 0.13 Sharpe amounts to almost nothing.

Every month, the rule refits the candidate laws on a rolling 1,610-observation window, divided into 1,127 training observations and 483 holdout observations. It repeats the 4,000-replication block bootstrap with block length five, then solves the lower envelope at a 5% annual reference. Holdings are c*q0, subject to a 10% per-name cap and 1.0 gross. Orders execute market-on-close. We charged 0.4 cents a share, imposed a $1 order minimum and modelled zero slippage. Uninvested capital earns point-in-time fed funds, with a fallback of 1.25% annual.

Sizing explains much of the Sharpe of 0.13 and Max Drawdown of -0.25%. At the paper's own 5%-reference optimum, the rule carries about 4.3% gross risky exposure against a budget of 1.0. Cash dominates the book by construction.

The paper reports no Sharpe, return series or drawdown for its own exercise. Total Return 0.10% and Sharpe 0.13 are our figures. The paper's self-reported results are prospect values: -0.04478 modelled and -0.04867 empirical at the 5% reference, followed by -0.09245 and -0.10049 at 10%. These quantities differ, so they permit no directional comparison between our results and the paper's. The author also describes the empirical pair as descriptive rather than out-of-sample, since the same 483 holdout days were used to define the ambiguity set.

Our run has two weaknesses. Its warm-up window contains 1,610 observations, about 6.4 years of daily data, against a 5.8-year traded window. The actual number of rebalances therefore depends on the pre-2020 history available. The direction we traded may also differ from the paper's long-short mean-variance ray. Neither weakness belongs to the paper.

A certificate for the discrete problem

The interval branch-and-bound is the paper's cleanest contribution. From 1,201 starting intervals and a 1e-10 tolerance, it certifies global optimality gaps of 9.20e-11, 1.00e-10 and 9.99e-11 at the three reference returns. The method avoids dependence on sign changes, which a root search over a piecewise-smooth CPT objective can miss at a tangency. The author states the boundary of the claim clearly: the certificate applies to the compressed 1,024-scenario discrete objective. It excludes scenario-compression and parameter-estimation error.

Most objections already appear in the discussion section. Uncertainty in q0 is ignored. The two-fund direction comes from a concave-utility result that CPT does not satisfy, leaving an optimum restricted to one ray. The ambiguity set consists of five fitted families. Proposition 2 does not carry over to their convex hull because the prospect functional is nonlinear in P. Turnover and transaction costs are absent. The paper says that a genuine evaluation would require an independent sample or a rolling design. Its universe remains fixed at 30 named tickers with 1,611 common trading days from 2 January 2020 through 1 June 2026.

The diagnostic deserves to survive. Table 8 shows the best-fitting mixture beating the multivariate Gaussian by 4.30438 on holdout log score. Even the two rejected mixtures finish 4.1 to 4.2 above it. Differences among the five survivors run from 1e-5 to 1e-2, with bootstrap intervals crossing zero. Between inverse gamma and the grid NPMLE, the dependence structure scarcely changes; the largest off-diagonal correlation difference is 0.0049.

The worst-case layer has yet to earn its price. In the only application shown, Table 10 places its exposure just 0.051 pp of gross weight from the point-estimate winner at the 5% reference and 0.080 pp away at 10%. We read the NPMLE's truncated grid as the reason it becomes active. A version in which a moment class or a Wasserstein ball stresses the ambiguity set, or in which the direction is re-estimated with its own uncertainty, could make the envelope bite. Along one ray, with five projected means packed between 0.99517 and 1.00000, it barely moves the decision.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each month-end rebalance date t at the close:
  1. Build a common-trading-day close matrix for the fixed 30-stock universe.
  2. Compute daily log returns in percentage units: X_t = 100 * log(close_t / close_{t-1}).
  3. Require 1,610 prior observations; otherwise skip the rebalance.
  4. Split the rolling window chronologically:
       - 1,127 observations for training
       - 483 observations for holdout/model comparison
  5. Estimate the common fund direction on training data:
       v = mean_train_returns - r_f * 1
       q0 = S_train^{-1} v / (v' S_train^{-1} v)
       c_max = max_leverage / ||q0||_1
  6. Fit candidate NMVM mixing laws and the Gaussian benchmark.
  7. Score models on the holdout set and retain models whose paired circular-block-bootstrap log-score difference interval contains zero.
  8. Generate compressed Sobol scenarios for each retained model.
  9. Solve for scalar exposure c in [0, c_max] that maximizes the lower envelope of CPT value across retained models, using the configured CPT gain/loss curvature, probability weighting, and loss-aversion parameters.
 10. Target weights = c * q0, clipped by max 10% per-name position size and gross leverage &lt;= 1.
 11. Execute rebalance at the same close as market-on-close execution; skip any trade lacking a real close price.