Thirty years of returns cannot reliably distinguish a long-only equity manager with 1.6% of true annual alpha from a lucky zero-alpha manager. Ulrich's calibration reaches that conclusion, and the arithmetic fits on the back of an allocation memo.
The detection budget
Two parameters characterize any domain. The first is sigma, the dispersion of an outcome conditional on the actor's true skill. The second, sigma_skill, is the cross-actor standard deviation of true skill. Their ratio gives the per-observation signal-to-noise ratio. Multiply that ratio by the square root of the effective number of independent outcomes and you have Ulrich's detection budget, a single dimensionless measure of what a career record can support.
Ordinary power analysis sets the threshold. At one-sided 5% significance and 80% power, the smallest detectable skill increment is about 2.5 sigma over the square root of N_eff. After normalization, the condition is SNR times sqrt(N_eff) greater than 2.5/d, with d denoting the skill gap in units of the skill distribution. A medium gap (d = 0.5) requires a budget above 5.0. For a large gap (d = 0.8), the requirement is 3.13. A comparison between two actors, instead of one actor against a benchmark, raises the constant by sqrt(2), to roughly 3.5.
Ulrich calibrates ten domains from published work (Wermers, Fama and French, Barras, Scaillet and Wermers, Cochrane, Korteweg and Sorensen, Bertrand and Schoar, Kaplan, Klebanov and Sorensen, Mellers et al., De Vany), then places them on a log-log map. The paper supplies no original dataset and no backtest. Budgets at the knowable end are large: about 250 for high-frequency systematic trading at a million trades, 33 for tournament chess and 16 for NBA career free throws. The figures then fall sharply. They are 5.6 for Good Judgment Project superforecasting, 2.4 for a venture partner, 2.0 for active equity, 0.98 for a film director, 0.58 for a non-founder Fortune 500 CEO, and effectively zero for a single-venture founder. Every mapped domain has SNR below 1. In all ten domains, per-observation noise exceeds the spread in skill, making detection an integration problem.
The basketball example sticks. A single championship lifts the posterior probability that a team is genuinely the best from a prior of 1/30 to about 20-25%. A best-of-seven favors the stronger team only about 71% of the time, and four such series must run in sequence.
Active equity gets 2.0
For long-only active management, Ulrich sets sigma between 0.03 and 0.08 per year and sigma_skill at about 0.02 per year. That gives SNR of about 0.40. Across a 30-year career, N_eff ranges between 10 and 40 after allowing for correlation of 0.05 to 0.15 among portfolio decisions. With N_eff of 25, the budget is about 2.0, short of the 3.13 needed for a large effect. Here, d of 0.8 corresponds to 1.6% of true annual alpha, which remains below the threshold.
Running the inequality backwards requires roughly 61 independent annual observations at SNR 0.40 to reach 3.13. The 61 comes from my arithmetic using his inequality; Ulrich does not print it. His own sensitivity table puts active equity at 2.19 even with zero correlation across 30 annual returns. Career length binds. Under the equicorrelation model, N_eff approaches 1/rho as N grows. At rho = 0.10, no career length produces more than ten.
Ulrich keeps the claim narrow. Top-decile managers with true alpha above 3% (d above 1.5) are described as marginally feasible. He immediately qualifies that opening through Denrell and Liu (2012): in noisy heavy-tailed domains, unreliable processes disproportionately generate extreme observed performance. Outcome inference is therefore weakest in the top tail. Each budget on the map also concerns a single actor chosen in advance. Selecting apparent stars ex post calls for the multiplicity correction he cites in Barras, Scaillet and Wermers (2010), which the map leaves unapplied.
Correlation carries much of the verdict
Ulrich's sensitivity table is the only part of the paper that stresses the verdicts against a parameter, and it is the paper's most useful exhibit. With SNR 0.40 and 30 annual returns, active equity scores 2.19 at rho = 0. The budget then drops to 1.40, 1.11 and 0.84 at rho of 0.05, 0.10 and 0.20. The verdict survives even at zero correlation, and Ulrich presents it that way.
Elsewhere, the same table shows how much depends on rho. The one boundary placement, superforecasting at 5.59, retains its budget because tournament design keeps cross-question correlation near zero. At SNR 0.25 over 500 questions, the budget falls to 1.10 once rho reaches 0.05. Ulrich consequently describes the placement as detectable-with-discipline rather than knowable.
The Fortune 500 CEO case is built two ways, with different levels. Table 1 uses 40 raw quarters at SNR 0.22 and produces 1.39 at rho = 0. The text instead constructs the placement from 5 to 10 truly independent strategic decisions during a tenure. At N = 7, the budget is 0.58. Ulrich asserts that count of independent decisions without deriving it, while the equicorrelation table supplies another path to the same unknowable verdict.
Ulrich gives the broader warning directly: the claims are supported by "the approximate region of the placement and not by the point estimates." The pooled cross-actor dispersion used in the binary calibrations also overstates sigma by a few percent. For the HFT figure, he concedes sigma_skill is "a rough approximation because the industry is famously secretive about per-trade economics." The placement still holds for any sigma_skill between 5 and 50 basis points because career N reaches into the millions.
What validates the process evidence?
Ulrich proposes replacing thin individual records with population-level validation of the kind used in medicine. One record cannot bear the scoring burden. A population of several thousand actors can estimate what an individual history cannot: which observable practices predict realized outcomes, and by how much. The evaluator then scores the individual on those practices.
Four worked examples support the prescription: active share for funds, Brier scores from the roughly 25,000-forecaster Good Judgment Project population, trait batteries validated across about 300 CEOs by Kaplan, Klebanov and Sorensen, and pharmaceutical stage-gate discipline. Finance supplies the contested one. Cremers and Petajisto (2009) found that funds above 80 active share outperformed, while closet indexers below 60 underperformed. Frazzini, Friedman and Pomorski (2016) argue that benchmark adjustment removes the predictive power, and Cremers (2017) responds. An allocator is left scoring managers on a feature whose predictive validity remains disputed in the journals. Ulrich declines to settle the exchange and uses the example "for its structure rather than as a settled empirical claim."
His defence carries more weight than the one-line prescription in the abstract. Process metrics, he argues, require stable prediction rather than causal identification. They should suffer less domination by noise than records in the unknowable region. He also treats process and outcomes as inputs to be integrated: the process metric provides a prior, the record provides a likelihood, and their reliabilities determine the weights. "In the essentially unknowable region the posterior is dominated by the process prior." Ulrich identifies the counters himself, including gaming, omitted variables, exchangeability with the validation population, and drift in the feature-outcome mapping across regimes. His sharpest structural limit follows from comparability. Validation favors generic practices, while a manager's idiosyncratic edge is unvalidatable by construction.
For an allocator, that limit consumes much of the prescription. The feature that can be scored across a fund population resembles the Bloom and Van Reenen layer of generic operational practice. The idiosyncratic part of the mandate, which is what the allocator is buying, has no population against which it can be validated. The surviving prescription is thinner: shrink the record heavily and treat any validated features as a prior rather than a verdict.
The other half of Ulrich's case survives intact, and that is where I would spend the effort. Sorensen (2007) attributes about two-thirds of apparent persistence in top-quartile VC returns to deal flow rather than investment skill. Ulrich draws out the operational consequence. A deal-flow advantage may belong to the franchise rather than the partners, so successions, departures and fund-size expansions may break it precisely when reinvestment decisions arise.
Paying for franchise gravity can be rational. It remains a different wager from paying for transferable skill, and the return series cannot reveal which one you bought.
We could not test any of this on our data. The framework replaces returns with manager process data, such as holdings-based practice panels and executive evaluations, and we have none of it. Reliability and shrinkage arithmetic could be applied to US equity and ETF strategy return streams, and we may do that. Such a run would test the general outcome-noise mechanism rather than Ulrich's mutual-fund or venture calibrations.
The figure I would take to an investment committee is the observation ceiling. In his equicorrelation model, rho = 0.10 caps a career at ten independent observations, while the inequality requires about 61 at SNR 0.40. A higher SNR than Ulrich assigns would change my view. That would mean sigma_skill above 0.02 per year against sigma of 0.03 to 0.08. Correlation offers no rescue because the active equity row already fails at rho = 0, with a budget of 2.19.