I would tune the shrinkage scalar before paying to train a covariance model on tracking error. Across the low-capacity models Yuan, Zhang and Su test, decision-focused learning (DFL) gains at most 1.76% over MSE. Their 16.5% equity tracking-error cut comes against a covariance-reconstruction target that the authors themselves warn against reading as an optimally specified forecaster. A matched neural comparison finds no aggregate DFL advantage in the tested architecture. The paper offers a useful account of why one-parameter models have so little room to gain, though the authors stress that rank one alone does not guarantee the result.
What does the covariance model buy?
The trade is a sparse index replica. From the top 100 S&P 500 names, reduced to N=97 by a coverage filter, the method selects K stocks, usually 20 by market cap. A long-only, fully invested quadratic program (QP) then chooses weights to minimize tracking variance against the index. The covariance matrix supplied to the QP is the sole learned input. Lower tracking error is the objective; the paper makes no alpha claim.
Training takes two forms. The two-stage predictor minimizes Frobenius loss against a covariance target. DFL instead puts the QP inside training, evaluates its weights on the next 21 days of squared tracking returns plus a turnover term, and differentiates through the QP's optimality conditions. Only the selected stocks' rows and columns enter that optimization. At N=100, K=20, they make up 36% of the matrix; MSE also fits the remaining 64%.
The authors trace the distinction through the predictor's Jacobian, the map from parameters to covariance entries. They call the limiting case Jacobian rank collapse. At rank one, every nonzero per-example gradient from either loss lies on a single line. Task training can reverse an update along that line, with no new direction available. A one-parameter Ledoit-Wolf style shrinkage model has rank one by construction. More strikingly, their 385-parameter conditional shrinkage network has rank one at every one of 54 measured inputs, with a median singular-value ratio around 1.2×10^-6. All its parameters feed one scalar intensity. Pointwise rank therefore need not grow with parameter count.
The equity data are daily returns from 2006 to 2025. Training rolls over 3 years, followed by 6-month validation and 1-year testing, for nine folds over 2016 to 2025. The authors extend the work to six markets, the S&P 1500 and synthetic shortest-path and knapsack problems.
Rank one has limits
Across 38 equity configurations, DFL's change relative to MSE ranges from −0.57% to +1.76%, with a 0.25% median. Sparsity shows no detected monotone association over K/N from 0.008 to 0.43 (ρ = −0.16, p = 0.35). At N=100, validation-tuned shrinkage, SPO+ and DFL pass a 20 bp equivalence test in all 12 pairwise comparisons. Each beats plain Ledoit-Wolf by 7 to 19%. Getting the scalar fitted matters more here than choosing among those three training objectives.
The authors treat that agreement as an empirical finding. Two per-example gradients can each occupy one line while their batch gradients point elsewhere: their two-parameter counterexample produces orthogonal batch updates from two rank-one examples. Synthetic results show what added capacity can accomplish. With full capacity, SPO+ lowers regret by 11.59% on shortest path and 10.59% on knapsack. Across eight comparisons, only knapsack survives Holm correction (adjusted p 0.0156); shortest path does not (adjusted p 0.0684). With one, two and eight update directions, changes span −0.61% to +1.89%.
A fresh-data follow-up across three batch orders finds full-capacity gains of 12.12/12.54% on path and 13.67/13.76% on knapsack at 20/80 epochs. Scalar gains remain below 0.6%. Then the authors keep the function class fixed and rescale its coordinates. At ε=0.01, the shortest-path gain falls from 11.79% to 1.19%; the knapsack gain falls from 10.17% to 0.52%. Adjusting the SGD step to compensate restores 11.79% and 10.17% exactly.
That rescaling experiment is the paper's most useful result for someone fitting these models. Plain SGD loses most of the gain when the coordinates change, even though the function class stays fixed. Conditioning matters here. Rank sets a ceiling, while a poorly scaled parameterization can fall well short of it. The spectral-gap bound meant to cover the intermediate case is vacuous at all 54 structured-model points, leaving little theoretical guidance between rank one and full rank.
The neural result and its control
At K=20, the headline neural result reports annualized TE of 0.0276 for DFL against 0.0331 for MSE: a 16.5% cut across 9/9 folds, with a 95% interval of [−26.1%, −7.1%]. Gains vary sharply by fold, from −0.3% in 2017 to 2018 to −40.0% in 2021 to 2022.
The MSE comparator learns to reconstruct trailing 63-day covariance. The authors explicitly warn that this should not be read as an optimally specified forecasting model. They therefore run a matched comparison using a 197-parameter residual network and one chronology for both losses. MSE targets future 21-day covariance; DFL uses those same 21 days for its task loss. Across three seeds and ten annual folds, mean TE is 5.711% for MSE and 5.715% for DFL. DFL is 0.07% worse on the mean despite winning six of ten years. Validation retains the untrained initializer in 17/30 MSE fits and 12/30 DFL fits. Their abstract reports no aggregate DFL advantage for this tested architecture.
The authors argue that architecture, support, targets and evaluation protocol all change between the historical run and the control, which therefore "cannot identify which change matters."
They are right.
The 40-epoch budget and narrow learning-rate grid add to the uncertainty. With validation retaining the initializer in 17/30 and 12/30 fits, the null result is weaker, and so is the case for attributing the 16.5% to DFL. The historical comparison measures DFL against a reconstruction target; it cannot establish a gain over a properly specified forecaster. In the one matched architecture tested, the mean gain is absent (−0.07%). Neither result persuades me to pay for DFL.
The scalar forward-target control reinforces that judgment. Fitting shrinkage to future covariance reduces covariance error by 37.25% yet increases tracking error by 4.13%. Selecting the scalar by tracking error on a 21-point validation grid cuts TE by 1.46%, the lowest mean TE among the estimators tested, compared with 0.69% for plain Ledoit-Wolf. The training target changes the trade.
Our traded book
We also built a version of the strategy. The figures in this section are ours, from a daily backtest covering January 2020 to July 2024. Our book holds the 20 largest names in an internally built, cap-weighted top-100 benchmark (non-ADR). It uses scalar shrinkage with one α per fold trained by DFL through the QP, caps each name at 10%, sets the turnover penalty to zero, and rebalances at the close every 21 trading days. We charged $0.004 a share with a $1 minimum.
Over that period, our traded DFL book returned 69.74% in total. Sharpe was 0.71, Sortino 0.88 and Calmar 0.36, with 22.03% annualized volatility and a −35.04% maximum drawdown. Those are absolute-return figures. The paper's tracking-error figures are 2.76% against 3.31% for its neural model, with scalar shrinkage gains under 1.8%. These measure different outcomes and cannot be compared. A fully invested long-only mega-cap basket can carry 22.03% volatility and a −35.04% drawdown regardless of whether DFL improves tracking against MSE.
We did not compute matched-date tracking error against an MSE-trained book, so we cannot say whether even the small scalar gap reproduces. Our run uses a forward 21-day MSE target, as in the authors' forward-target controls rather than their main equity runs. It also has a 10% name cap, which they test only as a robustness check, and our own benchmark construction. The paper's figures do not describe our book. This was one automated pass, and our quick automated test does not settle the paper's tracking-error claim either way.
Before spending the compute
The paper gives practitioners much of what they need to attempt replication, along with its own warnings. The scalar baseline tuned on validation tracking error is the place to start: at N=100, it is equivalent to DFL within 20 bp. For one parameter, DFL training took about 13 minutes per fold, roughly 3,329 times the cost of POET plus QP. POET is a factor-plus-thresholding covariance estimator; paired with the QP, this timed cheap baseline took 0.23 s per fold.
Turnover deserves the same attention as TE. At K=20, pure DFL traded 0.030 against 0.003 for MSE. Under a cost-aware QP at 100 bps, the reported net figure still favors DFL, 0.0280 against 0.0331, though that comparison retains the reconstruction target. BD-DFL adds a relative MSE penalty. At N=451, K=10 and β=0.001, it cut TE by 15.6%, against 13.0% for pure DFL; the authors describe the coefficient as an empirical choice. The best-tested rows at N=100 and N=478 are labelled exploratory. Their forward-target controls also use a frozen, survivorship-biased universe, as the authors flag.
Point-in-time membership and longer training for both losses are on the authors' own list of open items. A tracking-error gap that survives those tests would change my view. Until then, I would tune the shrinkage scalar on tracking error and spend the compute budget elsewhere.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.