AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

DFL's 16.5% tracking-error cut fails a matched control

Rank-one gains stay under 1.8%; a matched 197-parameter network puts DFL 0.07% behind MSE.

2026-10-01 · 7 min read · Portfolio Optimization · US equities

Reviewing: Jacobian Rank Collapse in Decision-Focused Learning · Aojie Yuan, Haiyue Zhang and Zijian Su · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

PIT-Capitalization Sparse Tracking: Forward MSE versus Decision-Focused Shrinkage

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (6)

  • Paper §3 alternative tracking score: s_i = w_i^idx (Σ̂_{i,:} w_idx)^2 / Σ̂_{ii}.: Use the paper's default index-weight/market-cap ranking, not tracking-score ranking. (invalidates: Paper results specifically concerning dynamic tracking-score selection and BD-DFL under that selection)
  • Paper Eq. (1): min_{w_S} w_S^T Σ̂_{SS} w_S - 2 w_S^T Σ̂_{S,:} w_idx, subject to 1^T w_S = 1, w_S >= 0.: Retain the exact objective and constraints and add the mandated w_i <= 0.10 cap. (invalidates: Paper numerical TE and turnover comparisons for its uncapped QP)
  • Paper §4.1 L_MSE = N^{-2} ||Σ̂(x_t; θ) - Σ_realized,t||_F^2, with a preceding-63-day target.: Keep N^{-2} and replace the historical reconstruction target with genuinely future-21-day sample covariance. (invalidates: Paper empirical comparisons that train MSE to reconstruct preceding-63-day covariance)
  • Paper Eq. (3) optional BD-DFL: L_task + β ||Σ̂(θ) - Σ_realized||_F^2 / (||Σ_realized||_F^2 + 10^{-12}); Σ_realized is trailing.: Do not activate optional BD-DFL; retain its formula as a reference without relabeling its trailing target. (invalidates: Paper empirical BD-DFL improvements)

2 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 69.7%Jul 2024
Sharpe
0.71
Total Return
69.7%
Max Drawdown
-35.0%
CAGR
12.5%
Volatility
22.0%
Beta vs SPY
0.99
Trades
2,158

What the paper reports for its own strategy

  • Neural DFL vs MSE, N=100, K=20, 9 annual test folds 2016–2025: annualized TE 0.0276 vs 0.0331 (−16.5%), one-sided Wilcoxon p=0.002, 9/9 folds. Main comparison is without transaction costs; turnover penalty γ in loss.
  • Same neural folds, K=20: worst-fold TE 0.0360 vs 0.0564 (−36.3%), max tracking drawdown 0.0306 vs 0.0362 (−15.5%), daily CVaR5% 0.0038 vs 0.0044 (−14.3%), TE std across folds 0.0045 vs 0.0098.
  • TC-aware QP, K=20: DFL net cost (TE + c·TO) 0.0280 vs MSE 0.0331 at c=100 bps. At K=10: 0.0466 vs 0.0511. DFL turnover 0.041 vs 0.001 at K=10.
  • BD-DFL (β=0.001), corrected N=451, dynamic selection, 11 folds: mean TE 0.0620 vs MSE 0.0734 (−15.6%), p=0.0024, 9/11 wins at K=10. At K=20: −5.5%, p=0.103.
  • Matched neural forward-target DFL, 2016–2025, no transaction costs: mean annualized TE 5.715% vs 5.711% for MSE (0.07% worse).

I would tune the shrinkage scalar before paying to train a covariance model on tracking error. Across the low-capacity models Yuan, Zhang and Su test, decision-focused learning (DFL) gains at most 1.76% over MSE. Their 16.5% equity tracking-error cut comes against a covariance-reconstruction target that the authors themselves warn against reading as an optimally specified forecaster. A matched neural comparison finds no aggregate DFL advantage in the tested architecture. The paper offers a useful account of why one-parameter models have so little room to gain, though the authors stress that rank one alone does not guarantee the result.

What does the covariance model buy?

The trade is a sparse index replica. From the top 100 S&P 500 names, reduced to N=97 by a coverage filter, the method selects K stocks, usually 20 by market cap. A long-only, fully invested quadratic program (QP) then chooses weights to minimize tracking variance against the index. The covariance matrix supplied to the QP is the sole learned input. Lower tracking error is the objective; the paper makes no alpha claim.

Training takes two forms. The two-stage predictor minimizes Frobenius loss against a covariance target. DFL instead puts the QP inside training, evaluates its weights on the next 21 days of squared tracking returns plus a turnover term, and differentiates through the QP's optimality conditions. Only the selected stocks' rows and columns enter that optimization. At N=100, K=20, they make up 36% of the matrix; MSE also fits the remaining 64%.

The authors trace the distinction through the predictor's Jacobian, the map from parameters to covariance entries. They call the limiting case Jacobian rank collapse. At rank one, every nonzero per-example gradient from either loss lies on a single line. Task training can reverse an update along that line, with no new direction available. A one-parameter Ledoit-Wolf style shrinkage model has rank one by construction. More strikingly, their 385-parameter conditional shrinkage network has rank one at every one of 54 measured inputs, with a median singular-value ratio around 1.2×10^-6. All its parameters feed one scalar intensity. Pointwise rank therefore need not grow with parameter count.

The equity data are daily returns from 2006 to 2025. Training rolls over 3 years, followed by 6-month validation and 1-year testing, for nine folds over 2016 to 2025. The authors extend the work to six markets, the S&P 1500 and synthetic shortest-path and knapsack problems.

Rank one has limits

Across 38 equity configurations, DFL's change relative to MSE ranges from −0.57% to +1.76%, with a 0.25% median. Sparsity shows no detected monotone association over K/N from 0.008 to 0.43 (ρ = −0.16, p = 0.35). At N=100, validation-tuned shrinkage, SPO+ and DFL pass a 20 bp equivalence test in all 12 pairwise comparisons. Each beats plain Ledoit-Wolf by 7 to 19%. Getting the scalar fitted matters more here than choosing among those three training objectives.

The authors treat that agreement as an empirical finding. Two per-example gradients can each occupy one line while their batch gradients point elsewhere: their two-parameter counterexample produces orthogonal batch updates from two rank-one examples. Synthetic results show what added capacity can accomplish. With full capacity, SPO+ lowers regret by 11.59% on shortest path and 10.59% on knapsack. Across eight comparisons, only knapsack survives Holm correction (adjusted p 0.0156); shortest path does not (adjusted p 0.0684). With one, two and eight update directions, changes span −0.61% to +1.89%.

A fresh-data follow-up across three batch orders finds full-capacity gains of 12.12/12.54% on path and 13.67/13.76% on knapsack at 20/80 epochs. Scalar gains remain below 0.6%. Then the authors keep the function class fixed and rescale its coordinates. At ε=0.01, the shortest-path gain falls from 11.79% to 1.19%; the knapsack gain falls from 10.17% to 0.52%. Adjusting the SGD step to compensate restores 11.79% and 10.17% exactly.

That rescaling experiment is the paper's most useful result for someone fitting these models. Plain SGD loses most of the gain when the coordinates change, even though the function class stays fixed. Conditioning matters here. Rank sets a ceiling, while a poorly scaled parameterization can fall well short of it. The spectral-gap bound meant to cover the intermediate case is vacuous at all 54 structured-model points, leaving little theoretical guidance between rank one and full rank.

The neural result and its control

At K=20, the headline neural result reports annualized TE of 0.0276 for DFL against 0.0331 for MSE: a 16.5% cut across 9/9 folds, with a 95% interval of [−26.1%, −7.1%]. Gains vary sharply by fold, from −0.3% in 2017 to 2018 to −40.0% in 2021 to 2022.

The MSE comparator learns to reconstruct trailing 63-day covariance. The authors explicitly warn that this should not be read as an optimally specified forecasting model. They therefore run a matched comparison using a 197-parameter residual network and one chronology for both losses. MSE targets future 21-day covariance; DFL uses those same 21 days for its task loss. Across three seeds and ten annual folds, mean TE is 5.711% for MSE and 5.715% for DFL. DFL is 0.07% worse on the mean despite winning six of ten years. Validation retains the untrained initializer in 17/30 MSE fits and 12/30 DFL fits. Their abstract reports no aggregate DFL advantage for this tested architecture.

The authors argue that architecture, support, targets and evaluation protocol all change between the historical run and the control, which therefore "cannot identify which change matters."

They are right.

The 40-epoch budget and narrow learning-rate grid add to the uncertainty. With validation retaining the initializer in 17/30 and 12/30 fits, the null result is weaker, and so is the case for attributing the 16.5% to DFL. The historical comparison measures DFL against a reconstruction target; it cannot establish a gain over a properly specified forecaster. In the one matched architecture tested, the mean gain is absent (−0.07%). Neither result persuades me to pay for DFL.

The scalar forward-target control reinforces that judgment. Fitting shrinkage to future covariance reduces covariance error by 37.25% yet increases tracking error by 4.13%. Selecting the scalar by tracking error on a 21-point validation grid cuts TE by 1.46%, the lowest mean TE among the estimators tested, compared with 0.69% for plain Ledoit-Wolf. The training target changes the trade.

Our traded book

We also built a version of the strategy. The figures in this section are ours, from a daily backtest covering January 2020 to July 2024. Our book holds the 20 largest names in an internally built, cap-weighted top-100 benchmark (non-ADR). It uses scalar shrinkage with one α per fold trained by DFL through the QP, caps each name at 10%, sets the turnover penalty to zero, and rebalances at the close every 21 trading days. We charged $0.004 a share with a $1 minimum.

Over that period, our traded DFL book returned 69.74% in total. Sharpe was 0.71, Sortino 0.88 and Calmar 0.36, with 22.03% annualized volatility and a −35.04% maximum drawdown. Those are absolute-return figures. The paper's tracking-error figures are 2.76% against 3.31% for its neural model, with scalar shrinkage gains under 1.8%. These measure different outcomes and cannot be compared. A fully invested long-only mega-cap basket can carry 22.03% volatility and a −35.04% drawdown regardless of whether DFL improves tracking against MSE.

We did not compute matched-date tracking error against an MSE-trained book, so we cannot say whether even the small scalar gap reproduces. Our run uses a forward 21-day MSE target, as in the authors' forward-target controls rather than their main equity runs. It also has a 10% name cap, which they test only as a robustness check, and our own benchmark construction. The paper's figures do not describe our book. This was one automated pass, and our quick automated test does not settle the paper's tracking-error claim either way.

Before spending the compute

The paper gives practitioners much of what they need to attempt replication, along with its own warnings. The scalar baseline tuned on validation tracking error is the place to start: at N=100, it is equivalent to DFL within 20 bp. For one parameter, DFL training took about 13 minutes per fold, roughly 3,329 times the cost of POET plus QP. POET is a factor-plus-thresholding covariance estimator; paired with the QP, this timed cheap baseline took 0.23 s per fold.

Turnover deserves the same attention as TE. At K=20, pure DFL traded 0.030 against 0.003 for MSE. Under a cost-aware QP at 100 bps, the reported net figure still favors DFL, 0.0280 against 0.0331, though that comparison retains the reconstruction target. BD-DFL adds a relative MSE penalty. At N=451, K=10 and β=0.001, it cut TE by 15.6%, against 13.0% for pure DFL; the authors describe the coefficient as an empirical choice. The best-tested rows at N=100 and N=478 are labelled exploratory. Their forward-target controls also use a frozen, survivorship-biased universe, as the authors flag.

Point-in-time membership and longer training for both losses are on the authors' own list of open items. A tracking-error gap that survives those tests would change my view. Until then, I would tune the shrinkage scalar on tracking error and spend the compute budget elsewhere.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each chronological annual test window:
  Fit using the preceding 36 training months; select using 6 validation months.
  Exclude training labels and validation holdings that cross a later window boundary.
  At each eligible annual update, verify screen availability and set fixed
    benchmark weights from the top 100 non-ADR stocks by capitalization.
  At each 21-trading-day rebalance, using information available by the prior close:
    Select the 20 largest current benchmark weights; use identical support
      for all branches.
    Estimate trailing covariance S and form Sigma(alpha) =
      (1-alpha) S + alpha * tr(S)/N * I.
    Fit one fold-level alpha by closed-form forward-21-day covariance MSE;
      fit the DFL alpha through the capped QP using subsequent 21-day
      tracking loss (gamma = 0). Keep alpha = 0 as an unfitted control.
    For each branch, solve the long-only, unit-sum tracking QP with
      each selected weight &lt;= 10%; skip incomplete panels or failed solves.
    Trade the DFL book at the rebalance close, liquidating and re-entering;
      hold shares, allowing weights to drift, until the next rebalance.