Where to damp
In the estimate, in the algorithm, or both.
The damped Schur complement can be applied in two places. The recursion can run on the raw estimate $\widehat\Sigma$ with each block conditioned on the others through $A-\gamma BD^{-1}B^\top$, or the estimate itself can be tapered, the cross-block scaled by $\sqrt\gamma$, and handed to a plain optimizer. With two blocks these are the same operation with the dial labelled differently. With more they are different paths between the same two ends. Either way they agree at both of those ends, and out of sample their best points are within about one percentage point of each other. The choice is decided by what has to be done with the result, not by performance.
The identity
Let the covariance be partitioned into a block and its complement, and let $\Sigma\circ C_s$ be the same matrix with every cross-block entry multiplied by $s$. The conditional covariance of the block under the tapered matrix is $A - s^2 BD^{-1}B^\top$, the damped complement at $\gamma=s^2$, and the companion vector is $\mathbf 1 - s\,BD^{-1}\mathbf 1$. So a minimum-variance portfolio of the tapered estimate is a member of the two-dial family of the 2024 paper, complement damped by $\gamma$ and companion by $\sqrt\gamma$. The recursion damps both by $\gamma$. With more than two blocks the taper damps every cross-block at once, while the flat bridge conditions each cluster on an undamped rest. These are different paths between the same two ends. The localization note proves the two-block case.
Minimum variance, out of sample
Thirty-two assets in four sectors, one factor per sector and a weak market factor, ninety days of returns, twenty-four random markets. Each row is a path from the heuristic end to the plain optimizer, and the entry is the true variance of the estimated portfolio in excess of the true optimum, in percent, along the dial.
| where the damping lives | dial 0 | best | dial 1 |
|---|---|---|---|
| in the algorithm: the flat bridge on the raw estimate | 19.5 | 18.5 at 0.2 | 55.9 |
| in the estimate: cross-block tapered by $\sqrt\gamma$, then minimum variance | 19.5 | 17.3 at 0.1 | 55.9 |
| in the estimate: cross-block tapered by $\gamma$, then minimum variance | 19.5 | 17.2 at 0.2 | 55.9 |
| both: tapered by $\sqrt\gamma$, then the bridge at $\gamma$ | 19.5 | 17.7 at 0.4 | 55.9 |
The ends coincide because at $0$ every path is the block-diagonal heuristic and at $1$ every
path is the plain optimizer on the raw estimate. The minima differ by a point and sit at
different places on the dial, which is what two paths between the same ends look like. A gap of several points between the
rows, or minima at different ends, would refute the claim that the placement is immaterial.
Script: experiments/where_to_damp.py.
The race, and the data setting the dial
For a Thurstone portfolio the damping has to live in the estimate, because a race is a functional of one joint law and a tapered covariance is one. That makes it the right place to ask whether the data can set the dial. Put a prior on the covariance whose mean keeps the within-sector blocks and shrinks the cross-sector blocks toward zero, choose the prior's strength by marginal likelihood, and race under the posterior mean. Twelve assets in three sectors, twenty random markets; the entry is the squared distance from the race under the true covariance, times $10^4$.
| estimator | 60 days | 250 days |
|---|---|---|
| race on the full estimated correlation | 62.1 | 13.3 |
| race with no correlation | 43.6 | 23.1 |
| cross-sector block set to zero | 46.2 | 11.4 |
| posterior mean, sector prior, strength by marginal likelihood | 47.0 | 11.3 |
| posterior mean, diagonal prior, strength by marginal likelihood | 46.6 | 12.6 |
| average of the race over 32 posterior draws, sector prior | 47.6 | 11.8 |
Three things follow. The sector prior beats the diagonal prior at 250 days, so the damping has
to be aimed at the cross-block; shrinking every entry, which is what ordinary shrinkage does,
is weaker. The marginal likelihood puts the sector prior where the best hand-set dial sits,
without a dial. And averaging the race over posterior draws adds nothing over the posterior
mean, for this loss; whether sampling helps a tail objective is open. A diagonal prior that
matched the sector prior at 250 days would refute the first claim.
Script: experiments/bayes_race.py.
The rule
Damp in the estimate when the result has to feed something else: a constrained optimizer, a tail objective, a race, or a data-driven choice of strength. Damp in the algorithm when the matrix cannot be inverted whole, because it is too large or rank-deficient, or when the structure is a tree or a set of knots rather than a flat block pattern, since the recursion never inverts the full matrix and its conditioning composes down a tree. Either way, aim the damping at the cross-block.