AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Special Markowitz binds mean and covariance shrinkage

The Stein-loss result is tight, while tradable calibration remains unresolved

2026-09-16 · 7 min read · US equities

Reviewing: Special Markowitz: Thermodynamic Formalism for the Joint Regularisation of Returns and Covariance · David Reinhardt · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Special Markowitz Joint Spectral Reliability Top-500 US Equities

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

This is not a replication of the paper

  • The reliability potential is not uniquely operationally specified: the paper allows several calibrations, including asymptotic spiked-model overlap, resampling stability, cross-validation, factor-model fit, and practitioner judgement. The implementation should therefore test a clearly specified rolling random-matrix/BBP calibration and robustness alternatives; this evaluates the selected calibration rather than a uniquely determined Special Markowitz implementation.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (6)

  • Remark 9 spike calibration psi_k=rho_k^2: rho_k^2 is estimated by preceding-data moving-block-bootstrap eigenvector overlap because the paper gives no finite-sample estimator. (invalidates: Any claim that the implemented psi_k equals the paper's asymptotic population overlap exactly)
  • paper bulk calibration psi_k -> 0 and Phi_k -> infinity: Finite implementation sets hard-bulk psi_k to 1e-6. (invalidates: Exact attainment of the reference-state limit for a finite calibrated bulk mode)
  • eq 24 unconstrained relative Markowitz solution: The closed form is retained only as a diagnostic; final weights are obtained by solving the required long-only, fully invested constrained quadratic program. (invalidates: Exact equality between final asset portfolio weights and the unconstrained closed-form modal portfolio)
  • Long-only, fully invested, 10% position cap: These empirical constraints are imposed by the task and backtest framework rather than by the paper's unconstrained spectral problem. (invalidates: Exact unconstrained modal portfolio weights as final holdings; exact final-portfolio reliability-gate ratios)

2 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 32.3%Jul 2024
Sharpe
0.30
Total Return
32.3%
Max Drawdown
-54.9%
CAGR
6.4%
Volatility
23.9%
Beta vs SPY
0.80
Trades
1,955

A single reliability weight per mode should govern both sides of the Markowitz estimate. Most desks instead tune two shrinkage knobs, using Ledoit-Wolf or a factor target for covariance and Black-Litterman or a flat haircut toward zero for means. Reinhardt argues that two modelling axioms, together with the loss function, determine one shared shrinkage law. He names the separated case Special Markowitz. Its modes decouple, allowing free energy and pressure to add mode by mode.

The paper supplies no method for calculating that per-mode weight. Remark 9 ends with a list of possible calibrations: asymptotic spike overlap, resampling stability, cross-validation, factor-model fit and practitioner judgment. We therefore chose one and tested it. Our calibration uses a moving-block bootstrap with 64 replications and 21-day blocks. We matched bootstrap eigenvectors to the empirical modes, then used squared overlap as the reliability score, with a floor at 1e-6. This overlap serves as a finite-sample proxy for the asymptotic quantity proposed in the paper.

On point-in-time top-500 US stocks, rebalanced monthly and net of commissions, the construction returned 32.34% in total from 2020-01-01 to 2024-07-01. Its Sharpe was 0.30, annualised volatility reached 23.94%, and maximum drawdown was -54.87%. The result is weak. A drawdown beyond half the book disqualifies it. These figures come from our calibration choice, while the paper reports no performance number of its own for comparison.

A shared weight for both estimates

Begin from a reference state worth believing, a prior mean vector and prior covariance written (μ_ref, Σ_ref). Whiten the sample covariance against that state,

A = Σ_ref^(-1/2) Σ̂ Σ_ref^(-1/2),

then take the orthogonal eigendecomposition A = UΛUᵀ. Two empirical quantities now belong to each eigendirection k. The relative variance is λ_k, with λ_k of 1 indicating agreement between sample and reference in that direction. The return component δ_k is the whitened mean difference projected onto the corresponding eigenvector.

A score is attached to every mode. The paper calls it information persistence, ψ_k ∈ (0,1], meaning the fraction of directional signal judged statistically real. Its negative logarithm gives the reliability potential, Φ_k = -ln ψ_k ∈ [0,∞). Full shrinkage is then

m_k = e^(-Φ_k) δ_k and λ_k - 1 = e^(-Φ_k)(λ_k - 1).

Theorem 16 names this the coupling identity. One scalar, e^(-Φ_k), retains the same fraction of return signal and covariance deviation from the prior. Both estimation problems therefore share the shrinkage intensity τ_k = 1 - e^(-Φ_k). As Φ_k approaches infinity, the mode returns to the reference: λ_k goes to 1 and m_k to 0.

The resulting portfolio rule is short. Under the unconstrained relative mean-variance problem with unit risk aversion,

w_k = δ_k/(λ_k + e^(Φ_k) - 1).

This can be read as Markowitz with a direction-specific ridge s_k = e^(Φ_k) - 1. Equivalently, the sample Markowitz weight is scaled by the gate R_k = λ_k/(λ_k + s_k) ∈ [0,1].

The proposed source of profit is simple: stop inverting noise.

The paper suggests a random-matrix calibration. For aspect ratio γ = n/T_obs, the Marchenko-Pastur bulk lies between λ_± = (1 ± √γ)². The paper treats sample eigenvectors inside that bulk as uninformative. Their ψ_k goes to 0, taking the associated mode back to the prior. Above the Baik-Ben Arous-Péché transition, a sample eigenvector keeps asymptotic squared overlap ρ_k² with its population counterpart. The proposed assignment is ψ_k:= ρ_k², hence Φ_k = -ln ρ_k². Spikes preserve their means and variances. The bulk, responsible for destabilising inverse covariance, is disabled in both legs together.

How far does the proof reach?

Two formal results do the heavy lifting. Proposition 7 uses the Cauchy functional equation. Suppose reliability multiplies when independent sources are composed, ψ_AB = ψ_A ψ_B, and suppose the desired coordinate is additive and vanishes at full reliability. Then Φ = -c ln ψ with c > 0 is the sole form available.

Proposition 12 carries more force. Consider C² losses whose σ-derivative separates additively, ∂σ D = a(λ) - b(σ), subject to D(λ,λ) = 0 and D > 0 otherwise. Also assume that the combined objective J{λ,s} has a unique minimiser for every λ > 0 and s ≥ 0. Within this class, the linear coupling

σ* - 1 = (λ-1)/(1+s)

holds if and only if D is the Stein loss,

c[σ/λ - ln(σ/λ) - 1]

with c > 0. Normalisation sets c = 1. The equivalence in both directions gives the result its value. I take it as support for the log-det objective over Frobenius distance. That comparison is mine. The paper does not discuss Frobenius or any quadratic covariance loss, excluding those alternatives through its restricted class rather than a direct argument against them.

Useful consequences follow from the remaining formalism. The optimal modal free energy has a closed form. For every λ_k > 0, each modal pressure is real-analytic in Φ_k. It is strictly convex for every non-trivial mode, meaning (λ_k, δ_k) ≠ (1,0). Cross-susceptibility disappears across distinct modes because ∂²P/∂Φ_j∂Φ_k vanishes for j ≠ k. Finite-dimensional Special Markowitz therefore has no intrinsic phase transition. Any discontinuity at the bulk edge enters through a hard calibration rule. Reinhardt states this plainly, with more care than the physics terminology might suggest.

Empirics lie outside the paper. It contains no dataset, backtest, Sharpe or turnover. Figure 1 is synthetic and described as inspired by an S&P 500 spectral density plot in Bouchaud and Potters. The author also declines two conclusions readers could otherwise infer. Both structural assumptions are presented as axioms: "Neither is derivable from first principles; each is a modelling axiom". An exact representation of pressure within Ruelle's formalism, including dynamics, invariant measure and entropy functional, remains open as well.

The scale-invariance axiom is the weak point

Principle 10 imposes η_k λ_k = s_k, making the ridge eigenvalue-independent by construction. Equal reliability scores produce equal fractional contractions toward the reference, regardless of how far the empirical deviations extend. Analytical nonlinear shrinkage instead varies shrinkage according to the eigenvalue's location. This contrast is my own. Proposition 30 compares Special Markowitz only with Ledoit-Wolf linear shrinkage, an isotropic Black-Litterman-like limit and constant-α ridge.

The paper's uniform limit shows where the distinction matters. Set Φ_k = Φ for all k. The covariance side becomes

λ*_k = (1-τ)λ_k + τ,

which is the Ledoit-Wolf linear form. The paper observes that linear shrinkage leaves returns untouched, whereas Special Markowitz shrinks them. Any edge over linear shrinkage must therefore come from variation in Φ_k across modes. Calibration determines precisely that variation, and calibration is unresolved.

The return side raises another concern. The mean difference δ_k is decomposed in the eigenvectors of A. Yet the paper labels bulk sample eigenvectors uninformative, even though those vectors supply the coordinates used to divide the mean. A direction assigned Φ_k near infinity may already have contaminated the projection of every other component. The choice of (μ_ref, Σ_ref) is also unspecified. That choice determines the fixed point completely.

What our run actually used

Our universe was the point-in-time top-500 US stocks by capitalisation, excluding ADRs. We used 756 trading days of daily adjusted closes, annualised at 252. The reference state took the cross-sectional mean of μ̂ and the diagonal of Σ̂, with a 1e-8 variance floor. Decisions were made monthly at the month-end close and executed at the next day's close, producing roughly 54 rebalances from 2020-01-01 to 2024-07-01.

The paper's problem (24) is unconstrained. We added long-only and fully invested conditions, plus a 10% position cap, because an implementable book requires them. Every fill incurred commissions of $0.0040 per share, subject to a $1.00 minimum and a 1.00% cap on trade value. All reported figures are net of those charges. Slippage is zero because execution uses the market-on-close print. Borrow, financing and market impact are unmodelled.

The 23.94% volatility and -54.87% drawdown belong to a long-only, fully invested 500-name portfolio with a 10% cap. Reliability scores came from the same 756 days used to estimate Σ̂, leaving the n modal parameters fitted in sample. Our calibration choice and diagonal reference state are the first suspects behind the 0.30 Sharpe.

One experiment could decide the practical question

Estimate free two-parameter shrinkage, assigning one intensity to the mean and another to the covariance. Then measure the out-of-sample cost of forcing both to equal 1 - e^(-Φ_k), with turnover matched. A Sharpe advantage of 0.2 or more for the free version, at matched turnover on the same universe, would end the coupling identity's case as a practical constraint. Until that experiment is run, the Stein-loss rigidity remains a standalone theoretical result and the portfolio rule remains unvalidated.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each month-end decision close:
1. Form the point-in-time universe from that calendar year's screening records:
   STOCK, non-ADR, ranked by capitalization, top 500.
2. Load the preceding 756 trading days of adjusted closes and compute simple returns.
3. Set optimizer capacity to zero for assets with less than 95% return coverage.
   Skip the rebalance if fewer than 10 assets remain eligible.
4. Estimate annualized arithmetic means μ_hat and unbiased sample covariance Σ_hat
   using a factor of 252.
5. Construct the reference state:
   μ_ref = 1 × cross-sectional mean(μ_hat)
   Σ_ref = diag(max(diag(Σ_hat), 1e-8)).
6. Whiten the empirical state relative to the reference state and compute a
   deterministic, descending symmetric eigendecomposition.
7. Run 64 moving-block-bootstrap replications with 21-day blocks. Globally match
   bootstrap eigenvectors to empirical modes and estimate squared overlap ψ_k,
   clipped to [1e-6, 1].
8. Convert reliability to Φ_k = -log(ψ_k), jointly shrink each modal mean and
   covariance-eigenvalue deviation, and reconstruct μ_SM and Σ_SM in asset coordinates.
9. Solve:
      maximize w'μ_SM - 0.5 w'Σ_SM w
      subject to sum(w) = 1, 0 ≤ w_i ≤ 0.10, leverage ≤ 4.0.
10. Submit target changes for execution at the next trading day's observed close.
    Skip any trade lacking an observed execution price; never synthesize a fill.
11. Hold until the next monthly rebalance and apply the recorded execution costs.

The research comparison uses otherwise identical sample-covariance Markowitz and covariance-only-shrinkage Markowitz variants.