Parallel books on the same signal cannot identify aggregate crowding when every book shares the same overlap. Vary their sizes and compare them on common dates, and the estimate captures each book's own sizing at the aggregate position already in place. Rodríguez Domínguez and Noguer i Alonso prove this result and put it in the abstract. With homogeneous overlap, an unrestricted date effect absorbs aggregate crowding exactly. No within-date estimator can recover it. Extra sleeves change nothing. Extra years change nothing. The paper states the problem best: "the comparison that makes the experiment robust is the one that prevents it from measuring the crowding capacity is about."
The mechanism and the proposed experiment
Edge equals an uncrowded mean less an erosion stock, which accumulates over time. Under the paper's primitives, the stock is scalar, linear in its own lag and the scale response, first-order Markov, stable, and has a steady state equal to that response. Those assumptions leave one dynamic form: h_{t+1} = a h_t + (1-a) c(W_t). Its kernel representation has weights summing to one.
Capacity is the deployment level where steady-state edge crosses an exit hurdle. A desk needs that crossing before adding capital. Existing proxies for deployed arbitrage capital, including short interest, return comovement, institutional holdings and their liquidity footprint, and fund flows, rely on different primitives. They have rarely been benchmarked against one another.
The paper proposes a switchback experiment. Parallel implementations of one strategy, called sleeves, receive randomly assigned deployment scales from a finite arm set. An assignment remains in place for L consecutive periods, forming a block, and treated sleeves are compared with controls on the same calendar dates. The authors derive the attenuation generated by each sampling rule, along with minimum-variance deattenuation weights. They also obtain the sharp identified set for capacity on a finite grid, a reporting rule with coverage, sample-size formulas for contemporaneous and staggered assignment, and the forgone-edge cost of running the experiment.
The calibration uses thirteen long-short factor-adjusted strategies formed from six Fama-French portfolio sets. The monthly sample covers 271 months from 2001:01 to 2023:07. Eight drivers are screened from 125 candidate series for the adjustment. Monthly residual volatility is 3.611 pp, while mean pairwise residual correlation is 0.186.
Persistence comes from an AR(1) fitted to 337 months of a detrended log borrow tail. The series is reconstructed from option-implied financing spreads over 1996-2024, producing a = 0.9177. Post-publication decline, net of a pseudo-discovery benchmark, supplies c(1) = 0.123 pp per month. The authors run no deployment experiment and say so in the roadmap to Section 4. Every numerical result comes from this calibration plus Monte Carlo on a panel built to match it.
The simulations follow the algebra
The fit is close. Under contemporaneous assignment, Var = 4 sigma^2 (1 - rho-bar)/P. Simulation matches that expression to three decimals for P from 10 to 400 over 6,000 periods. Staggered assignment reaches a floor at 2 rho-bar = 0.372, after which more sleeves cease to help. The simulated variance ratio between the schemes is 1.6 with ten sleeves, 12.1 with a hundred, and 48.2 with four hundred.
The appendix supplies the full grids: 500 points on [0.05, 0.995] for fitting the geometric kernel, a pre-fixed set of 41 equally spaced arms, a budget of 120,000 sleeve-periods, and replication counts for each table. The headline timetable then follows directly. With 100 sleeves and two-year holds, 80% power requires 509 months. With 400 sleeves and four-year holds, it takes 76 months. Even before attenuation, distinguishing a 12 bp erosion from 3.6% monthly residual volatility consumes 17,993 strategy-months.
Everything contentious comes after that arithmetic.
A two-year hold captures 59%
A two-year hold captures 59% of the erosion. More precisely, block-average sampling recovers G_L of the steady-state slope: 0.082 for a one-month hold, 0.402 at twelve months, 0.595 at twenty-four, and 0.771 at forty-eight. Period-by-period randomisation identifies under a tenth of the target. The paper correctly describes this as an estimand failure rather than an inference failure.
Recovering steady state requires division by G_L, which requires a. That persistence parameter is transported from another market. If the true value is 0.946, the 509 months rise to 841. At 0.97, the requirement becomes 1,981. At 0.80, it falls to 259. A geometric deattenuation kernel applied to a true twelve-month finite-memory process biases the estimate by +30%. Under a longer tail at 0.97, the bias is -50%.
The authors identify this bridge themselves. Section 4.4 calls transported persistence "what a study should try hardest to pin down", while the conclusion argues that declaring the bridge is cheaper than removing it. Their own cost estimates make the weakness harder to dismiss. A hundred sleeves held for two years need eight years of calendar time to place a in [0.39,0.99]. Sixteen years narrows the interval to [0.66,0.99]. Sixty-four years reaches [0.82,0.98], yet the recovered effect remains uncertain by a factor ranging from 0.74 to 2.54.
Learning the kernel within the experiment looks like the natural repair. The paper's results advise against it. The feasible deattenuator has bias -0.537 and s.d. 0.540 for a target of -0.123. Even at a hundredfold sample, it remains the worst of four estimators. Given equal budgets and geometric truth, transporting the wrong kernel beats estimating it, with RMSE 0.003 against 0.058. The same ranking holds under two speeds, 0.008 against 0.032. Finite memory reverses it, at 0.037 against 0.014.
Panel-internal persistence identification gives q-hat = 0.28 for the four non-overlapping momentum strategies and 0.27 for all thirteen. The moving-block bootstrap interval is [0.00, 0.99]. As the authors state, that interval contains the transported 0.9177. It supports neither side of the bridge.
One kernel-free number remains available to an implementer. The terminal contrast gives a lower bound on the steady-state effect without any kernel assumption. It should be reported. The thirteen per cent shortfall at a two-year hold, used to make the bound appear tight, is calculated from the calibrated persistence. The bound itself is model-free. Its advertised sharpness depends on the calibration.
Arm placement decides the invoice
This part belongs in front of a risk committee. A zero arm compared with a unit arm needs 509 months and sacrifices 228 cumulative percentage points of edge. Arms at 1.5 and 2.0 require 501 months while giving up 11. Expanding the pair to {1.5, 3.0} cuts the experiment to 39 months at a cost of 13 points.
Net edge is flat near the scale a book would ordinarily run. An arm near that point costs second order in its distance but adds first order to the identifying gap. The paper derives an exact floor of 10.6 cumulative points for every symmetric pair around the optimum. Getting below that floor requires wide arms. The {0,4} pair destroys edge at 2.05 times the rate at which optimal deployment creates it, for 10 cumulative points overall. The paper's prose reports 9.8.
The optimum is 1.73, and own-sleeve capacity is 2.675. Both values depend on curvature of 0.05 in the scale response and an exit hurdle of 0.150. The authors label those inputs as design choices rather than estimates. Cheap arm placement therefore depends on the parameter the experiment is meant to learn. We have raised the same objection about a tuning constant selected with hindsight in our note on Wasserstein-robust allocation. This paper handles the issue better by declaring the choice in a table instead of fitting it quietly. The choice still sets the bill.
Detection and resolution carry different prices. A bracket formed from estimated arm means covers true capacity 0.479 of the time on a fixed grid, versus a nominal 0.90, and 0.514 under adaptive refinement. The valid alternative is a Bonferroni band over the entire candidate set. Its coverage is 1.000, with deployment-unit lengths of 2.386 and 2.774. Those lengths equal 7.12 and 8.27 multiples of the resolution scale of 0.335. The range has a ceiling of 4 and a truth of 2.675, so the resulting set says almost nothing. Adaptive placement produces the longer band. An arm whose band straddles the hurdle can never serve as an endpoint, which means concentrating arms near the crossing spends them where the band cannot use them.
Can aggregate crowding be recovered?
The abstract offers two routes around the impossibility result. One gives sleeves deliberately different exposure to the common position. The other uses time variation in that position.
The paper prices the first route for one configuration. Six sleeves divided into two disjoint blocks produce a norm of 0.27 for the component of aggregate crowding orthogonal to the calendar direction. Homogeneous overlap gives exactly zero. This route also requires an estimated overlap matrix. We did not find an estimator for that matrix based on holdings or borrow data. The second route compares across dates, bringing back the common component and multiplying the calendar requirement by 23.9 at a hundred sleeves.
The experiment therefore delivers own-sleeve capacity at a fixed average. When aggregate crowding matters, the authors say the design understates erosion and reports capacity too high. Standard desk impact models err in the same direction and by more. If crowding accounts for 15%, 30% and 50% of total erosion, an execution-cost model estimates capacity at 2.950, 3.311 and 4.030, compared with own-sleeve capacity of 2.675. The corresponding overstatements are 10%, 24% and 51%. The wedge exists, and the experiment measures it. Neither estimate brackets strategy capacity from both sides.
We could not run this design on our own data. The reason follows directly from the paper's estimand: deployment must be assigned. The experiment needs parallel books whose sizes were randomly set, held in blocks, and recorded at realised deployment for each book. Price histories contain none of those assignments. The erosion stock being deattenuated would also have to come from capital we deployed, which it did not.
A short multi-horizon pilot, powered for kernel shape rather than level, would change my view. The paper recommends exactly that approach, with no transported persistence afterward. Its budget exercise shows that finer cohort splits surrender more precision than they gain. RMSE rises through 0.056, 0.086, 0.105, 0.136, 0.162 as the design moves from two to eight cohorts. Any pilot must therefore stay coarse.
Until somebody runs it, the identification results stand, while the 42-year headline remains one priced scenario among four: 259 months at a = 0.80, 509 at the transported 0.9177, 841 at 0.946, and 1,981 at 0.97.