At 63 daily observations and 500 stocks, the sample covariance matrix's third principal component averaged about 50 degrees from the direction it was meant to estimate. The first missed by 17 degrees and the second by 39. These averages come from 1,000 simulated paths calibrated to resemble the US equity market. A larger stock universe does not cure the error.
Risk-model users should remember the 50 degrees. Bernstein, Goldberg, Gunther, Kercheval, Lan, Lin and Yao quantify the error in the standard construction under conditions familiar to quantitative equity: thousands of names and only a few months to a year of usable history.
The decomposition
The model is the familiar latent setup. Returns obey y = Bf + z, where the loading matrix is deterministic, f contains k factor returns, and z is idiosyncratic noise. The targets are the principal directions, meaning the orthonormal eigenvectors of the systematic covariance BΣ_f B'. They are unit-length exposure combinations ordered by explained variance. The leading eigenvectors h_j of the scaled sample covariance YY'/(np) estimate them. The asymptotic regime has high dimension and little history: p grows without bound, n remains fixed, and the analysis conditions on the realized factor path.
Theorem 1 gives almost-sure limits for two components of the squared sine of the angle between target and estimate. The first, out-of-subspace error, is δ²/(nλ_j + δ²). It measures the distance between the sample eigenvector and the true exposure span. Idiosyncratic noise forces the principal component outside that span, with an error that persists. The other component is rotation inside the span. It appears because the estimate has only n factor-return draws, leaving the realized factor second moment different from its population counterpart. Using the realized signal-to-noise ratio SNR_j = nλ_j/δ², the limit becomes 1/(1+SNR_j) + [SNR_j/(1+SNR_j)] sin²∠(ν_j, e_j).
The limit remains finite and positive. Expanding the cross-section neither magnifies the error nor drives it away. As the authors put it, a larger cross-section provides no additional draws of f. More history is what helps. With p = 3000, roughly the Russell 3000, increasing n from 20 to 250 days reduced average total error from 30 to 8.5 degrees for factor 1, from 56 to 19 for factor 2, and from 66 to 24 for factor 3.
Even after a full year of daily observations, the third direction remains 24 degrees off.
A floor visible in the eigenvalues
Theorem 2 has direct operational value. Form the n × n dual matrix W = Y'Y/(np), whose nonzero eigenvalues match those of the large sample covariance matrix. Write its j-th eigenvalue as θ_j, and let ℓ denote the average of the trailing n - k eigenvalues. The ratio ℓ/θ_j converges almost surely to the same limit as the out-of-subspace error. Everything needed to calculate it is already in the data. It supplies an asymptotic floor: Theorem 2's inequality compares p → ∞ limits, making limiting total error at least the limit of ℓ/θ_j. Equality requires the rotation term to vanish.
Most of the simulated error lies in this floor. At p = 3000, its share of predicted error in sine-squared units ranged from 80.2% to 92.7% across all three factors and every n between 20 and 250 days. Factor 1 moved from 88.5% at n = 20 to 91.7% at n = 250. Factor 3 began at 92.7%, finished at 90.4%, and reached 88.2% at n = 90. The larger component is therefore the one that can be estimated.
A companion result extends the argument to the full subspace. The sample span of the top k eigenvectors is an inconsistent estimator of the loading span because ½‖Π_H − Π_B‖²_F converges to a strictly positive sum of the floors. Yet summing ℓ/θ_j consistently estimates the distance between the subspaces. The span remains unrecoverable, while its separation from the estimate can be priced.
What returns leave unresolved
Theorem 3 goes beyond inconsistency. Hold fixed the data Y, the loadings, the realized path F, the particular returns Z, the constant δ², and the eigenvalues of K. Then allow the factor covariance to vary across matrices compatible with that spectrum. When k ≥ 2, rotation error spans every value in [0,1]; total limiting error ranges from the out-of-subspace floor through 1. Without further distributional assumptions about the factor path, the data impose no restriction on rotation error. Adding the factor path itself to the observables does not alter that conclusion.
The paper also gives a remedy. Corollary 3 makes rotation error vanish in the large-n limit. With less history, the user needs an exogenous estimate of G_B and Σ_f, along with a prior on FF'/n. If G_B, F and Σ_f are all known, Theorem 1 yields a computable, consistent estimator of the complete error.
Readers of the James-Stein eigenvector literature cited by the authors should keep one distinction in view. At k = 1, rotation error vanishes identically. The paper says this recovers the case of Goldberg et al. (2022) and the boundary result of Jung et al. (2012). Non-identification begins at k ≥ 2, the relevant setting for any multi-factor PCA model.
Conditions in an equity panel
Using the floor in practice depends on its assumptions surviving an equity cross-section. Four conditions matter. The authors identify two of them, cross-sectional independence and constant average specific variance, without pursuing either relaxation.
- The factor count
kmust be known. The paper observes thatℓandθ_jare computable oncekis given, while the simulation imposesk = 3instead of estimating it. A misspecifiedkputs the wrong eigenvalues into the average definingℓ. - Conditional on the factor path, specific returns must be independent across the cross-section. The authors conjecture that weak dependence could replace this condition and explicitly leave the extension unpursued.
- Cross-sectional average specific variance
δ²must converge to the same level on each date. Assumption 3 imposes this requirement, which the authors say may be relaxed without doing so here. The paper separately warns that its conclusions may fail when the history includes a shock shared by factor returns and specific variances. Assumption 2 excludes such a regime shift. March 2020 fits that description exactly. - The
p-limit must approximate realistic dimensions reasonably well. In the authors' pathwise plots atp = 3000andn = 250, the gap between finite-perror and the asymptotic limit narrows steadily withn, although they describe the remaining difference as non-negligible. They argue that thep-asymptotic limit still approximates true finite-n, finite-perror considerably better than zero, which is then-asymptotic limit. They interpret this as evidence for the HL regime at these values. How accurate the limit is atp = 3000remains unquantified.
The authors add two caveats. Every experiment holds one draw of the loading matrix fixed. They note that this choice leaves the asymptotics unchanged while affecting small-p results. The finite-p estimator ℓ/θ_j also tended to underestimate true out-of-subspace error in the simulation, and the paper says it has not investigated how generic that behavior is.
Theory and experiment cover different ground. Serial dependence and heteroskedasticity in factor returns are allowed by the theory. In simulation, factor returns are iid t(6), while specific returns are iid t(5). Annualized factor volatilities are 0.16, 0.08 and 0.06, compared with 0.40 for specific returns. The awkward empirical cases remain permitted by the theory and absent from the experiments.
Our implementation
We used the diagnostic as a risk control, extending the paper rather than implementing a proposal made by its authors. At each rebalance, the universe begins with the annual top 1,000 US stocks by capitalization, excluding ADRs. Eligibility requires complete closing prices for a short estimation window of 21, 63, 126 or 252 days, plus a strictly earlier 252-day anchor window. Dates with fewer than 500 qualifying names are skipped.
From the short window, we extract the top three PCA directions and calculate q_j = ℓ/θ_j. Each direction is pulled toward its projection on the anchor subspace using alpha_j = q_j, after which the directions are orthonormalized. The reconstructed covariance enters a long-only, fully invested minimum-variance optimizer subject to a 10% position cap. Weights use information available through the prior close, with execution at the same-day close.
All figures accompanying this article are ours. They span 2020-01-01 to 2024-07-01 using daily bars and include commissions of $0.0040 per share, subject to a $1.00 order minimum and a cap of 1% of trade value. Market impact, pass-through fees and taxes are unmodelled. Borrow and financing are irrelevant for this long-only, fully invested book. The same construction makes the reported volatility of 13.25% a property of a fully invested long equity mandate. It should be interpreted in that setting, rather than as a verdict on the estimator.
The run produced a 30.66% total return and a Sharpe of 0.52 from 2020-01-01 to 2024-07-01. The paper supplies no portfolio return, Sharpe or drawdown, leaving no like-for-like author result for comparison. Our result comes from one automated pass based on the paper's description, and alpha_j = q_j is our heuristic.
Two choices could plausibly outweigh any effect from the theorem. We imported k = 3 from the simulation without claiming that three is the correct count for the 2020s US cross-section. The test period also includes March 2020 and 2022, exactly the shared shock to factor returns and specific variances excluded by Assumption 2. The max drawdown of -30.98% occurs within that period. This weak result speaks to our implementation. We previously reviewed an optimizer study that treated factor inputs as exact (our note); the present paper shows why those inputs may miss by tens of degrees.
The diagnostic's proper use
The decomposition is exact for every p, and the eigenvalues of the dual matrix make the floor computable. An error of 50 degrees at n = 63, p = 500 should change how traders read factor attribution. Returns alone reveal none of the remaining rotation error. Recovering it requires exogenous G_B, Σ_f and a prior on FF'/n. Claims of factor-level precision from a short-window PCA model therefore present a lower bound as though it were a measurement.
I would treat the floor as a calibrated number after two further checks. The first is an empirical study on a real panel that estimates k instead of imposing it and includes a volatility regime change within the window. The second is evidence on whether the underestimation observed by the authors in simulation persists in that panel.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.