NCO discards a specific piece of the global minimum-variance solution, and Cotton restores an approximation by adding one asset per cluster. That trade is the paper. The method never requires a linear solve larger than max(k, largest cluster size), matching what NCO already pays. No data enter the result.
What NCO leaves behind
NCO first partitions the book. It solves minimum variance within each cluster using that cluster's covariance block, then repeats the exercise across the cluster portfolios. Stability comes from the outer problem having one dimension per cluster. Each cluster can respond to the rest of the book only through its budget, a single scalar.
The unconstrained global minimum-variance portfolio has the same two-tier form under block inversion. Slice Σ⁻¹1 by cluster. Each slice is (Σᶜᵢᵢ)⁻¹bᵢ, where the cluster block has become its Schur complement against every other asset and the companion vector bᵢ replaces NCO's vector of ones. Conditioning carries all the cross-cluster information.
Cotton reads NCO as this same construction after removing the conditioning. Putting it back requires, for every cluster, an inverse whose size is n minus the cluster. A two-stage method exists precisely to avoid that calculation.
A whole cluster represented by one asset
Cotton borrows the truncation from spatial statistics. Choose one member from every cluster and call it the knot. Cluster i is then conditioned on the other clusters' knots rather than on all outside assets, shrinking the inverse to (k−1)×(k−1).
Under Cotton's gateway model, this truncation loses nothing. After a cluster's non-knot members are regressed on their own knot, their residuals are uncorrelated with every asset outside the cluster. The resulting cross-cluster covariance blocks have rank at most one.
Cotton describes the price plainly: "a strong assumption, and clustering alone does not imply it". Residuals from members can be orthogonal to every knot while remaining correlated with one another across clusters. Regressing members on their own knot and the other knots therefore does not test the model.
The omitted correction Δᵢ measures what is lost. It is the part of cluster i explained by outside assets beyond their knots, and calculating it requires the large inverse. Cotton acknowledges this, treating Δᵢ as a one-time diagnostic rather than a component of the allocation.
A damping parameter γ in [0,1] creates the bridge. At γ=0, the construction is NCO with minimum variance at both tiers. At γ=1, the model gives Σ⁻¹1.
Frozen cluster interiors
Under the gateway model, member weights within a cluster are proportional to Eᵢ⁻¹(1 − βᵢ). They are independent of γ and every other cluster. NCO and the global minimum-variance portfolio therefore own the same positions inside each cluster, apart from the scale assigned to that cluster.
Only k knots are repriced along the bridge. Consider k knots with unit variance and common positive correlation ρ, using budget vector u=1. Knot exposure starts at 1 at the NCO endpoint and declines to 1/(1+(k−1)ρ) when γ=1.
Two consequences matter. First, the effective damping varies by cluster:
λᵢ/(1−λᵢ) = rᵢ · γ/(1−γ),
where rᵢ is the fraction of knot i's variance left unexplained by the other knots. When those knots nearly explain it, the exposure moves little until γ approaches 1. A shared γ does not produce a shared adjustment.
Second, take any cluster portfolios whose holdings remain within their respective clusters. Their outer covariance equals KΣ_PP K plus a diagonal matrix of member residual variances. In other words, the covariance of the knots is sandwiched between the clusters' knot betas, with an explicit ridge added. Cotton presents this as one explanation for the stability of NCO's outer tier.
Where does the optimum fall?
Cotton next introduces estimation error and asks where γ should sit. In an example with two clusters, knot correlation c=1/4 and δ=1, the estimate is 0 or 1/2 with equal probability. Exact arithmetic gives a unique global optimum of γ*=2/3. The gains are F(0)−F(2/3)=1/576 and F(1)−F(2/3)=1/900. NCO is the weaker endpoint in this case.
The four-asset quartic example reverses the endpoint ranking. Its optimum is γ*≈0.5587, while F(0)−F(1/2)=3/1210 and F(1)−F(1/2)=5/1452. Here full coupling is worse.
The ranking of the two ends flips between two examples in the same paper.
Cotton's local theorem examines γ=1. Under the gateway model, expected out-of-sample variance expands as V₀ + τ²G, and the optimum shifts to 1 − [G′(1)/V₀″(1)]τ². One derivative supplies the sign, and examples produce both signs.
For the three-cluster asymmetric-noise example, member variances are 1, 4 and 8. It yields G′(1)≈−0.0058498 and V₀″(1)≈2.3949, making full coupling a strict local minimizer for every sufficiently small τ. Exact checks at τ=1/10 and 1/100 find F(γ,τ)>F(1,τ) for γ in {0.999, 0.99, 0.9, 0.5, 0}. Cotton says these checks are consistent with a global optimum at that endpoint, though they do not prove it.
The symmetric k=10 family uses a true knot correlation of 1/4. Under small noise, the interior optimum wins when δ<8, while full coupling wins when δ>8. Each cluster reduces to an independent asset with variance 1/δ. Higher-order terms settle δ=8.
A fourth-order prize
The potential gain is small. Movement away from γ=1 is second order in the noise. If G′(1)>0, the variance reduction is fourth order:
G′(1)²/(2V₀″(1))·τ⁴.
At τ=0.05 in the quartic example, the predicted optimum is 0.9894 and the computed value is 0.9896. Cotton makes the limitation explicit: "A large gain at finite noise should not be inferred from the local sign." His δ=12 example approaches the issue from the other direction. It remains at γ=1 under small noise, then moves to γ*=10/11 when τ reaches 1/4.
The assumed noise is an additive symmetric perturbation of Σ. The partition and knots stay fixed while τ changes, a condition Cotton says is required for smoothness. Finite-history sampling error behaves differently and can also change the clustering.
Duplicated knots break the path
With two identical knots, the γ parameterization fails. The global optimum has variance 1/3. For every γ<1, the bridge produces 3/8. At γ=1, per-cluster pseudoinverses produce 1/2 because each cluster throws away a shared return still needed by the other cluster.
The limits in γ and in a ridge ε do not commute. Cotton's repair uses another path, parameterized by effective exposure λ. It is continuous, reaches NCO at λ=0 and minimum variance at λ=1. He explicitly declines to claim improved out-of-sample performance.
The missing US equities test
Every result above is analytic. The examples use covariances of four to six assets, alongside a symmetric ten-cluster family reduced to one scalar per cluster. Cotton leaves "an empirical comparison" for future work. Since the paper contains no real-data measurement, we are building the bridge on US equities ourselves.
Two quantities will decide whether it survives contact with that market. The first is Δᵢ on real sample covariances after clustering. Exactness at γ=1 depends on the gateway model. If the omitted correction is large, the right endpoint becomes an approximation with error Δᵢ, and nobody has measured that error.
The second is γ estimated from trailing data. Every example chooses γ against the population covariance. We found no out-of-sample estimation procedure for it in the paper.
Trading costs bear on both questions. Relative weights within each cluster remain fixed, so the bridge changes k knot exposures and k cluster budgets. Every name still trades whenever its cluster budget changes. Long-only presents the harder problem. Cotton states that box or long-only constraints eliminate any exact block decomposition of the optimum, both here and on the tree.
Cotton closes with the central qualification: "Neither end owns the bridge. An interior optimum is not a consequence of the architecture, and a statement that it is generic needs a noise model." He presents the contribution as a local theorem at the minimum-variance endpoint, supported by exact examples that go in both directions. He does not present it as guidance for choosing γ.
The objection remains. A bridge whose optimum is determined by the assumed noise model still lacks an allocation rule.
The structural result stands on its own: NCO is the zero-coupling endpoint of a continuous construction, rather than a separate philosophy. The allocation result remains a tuning parameter looking for calibration, much as we concluded about Stein-loss shrinkage in an earlier note. A single real covariance with Δᵢ small enough for γ=1 to mean what it claims would change my view.