The bond MDP's return edge deserves a discount until the policy is scored beyond the chain it was built on. Ramachandran, Iyer and Jain's more useful finding for a desk is how much a discretized yield chain can lose through truncation.

Where the return comes from

The authors use three default-free zero-coupon bonds, with 2, 5 and 10 year maturities, and rebalance monthly over six or twelve months. A Dynamic Nelson-Siegel curve supplies yields through its level, slope and curvature factors. Each factor follows an independent AR(1): the VAR matrix and shock covariance are both diagonal. The parameters come from Pericoli and Taboga's Markov-switching model, although the solved model has a single regime, as the authors state.

They simulate the VAR, snap the three yields to a 40bp grid and count transitions to form a time-varying transition matrix. Each state retains its 25 most likely successors, with probabilities renormalized. A cap of 80 states per date leaves 341 states in the six-month problem. The policy chooses among six fixed allocations: the three single-bond corners, equal weight, (0.2, 0.3, 0.5) and (0.5, 0.3, 0.2). It maximizes risk-neutral expected terminal wealth. Trading incurs a proportional cost c on the absolute change from drifted weights. Backward induction takes about 29 seconds per cost level in pure Python.

Mean reversion drives the result in our reading. The factors revert, allowing the current curve to forecast the next move, and the policy shifts duration accordingly. Its benchmarks stay put: an equal-weight ladder held to the horizon and a static mean-variance portfolio whose moments come from separate simulated paths. The authors disclose that the mean-variance choice is a single-bond corner in every scenario. They describe it as "the single best maturity given the initial curve".

At c = 0, six-month mean terminal wealth reaches 1.0860 for the policy, versus 1.0463 for static mean-variance and 1.0431 for the ladder. A separate run on the same curve (4,000 paths, three seeds) gives the 4.00-point gap a paired t of 86. Over twelve months, the policy reaches 1.1801 against 1.0901 for mean-variance.

What happens to the tails?

The authors acknowledge the trade-off: "the gain from dynamic allocation is a return gain, not a risk reduction." At c = 0, the policy's six-month fifth percentile is 0.9500. Mean-variance reaches 0.9616 and the ladder 0.9534. Mean drawdown tells a different story: 3.92% for the policy, level with mean-variance at 3.93% and below the ladder's 4.40%. The fifth percentile deteriorates even though drawdown does not. That fits the risk-neutral objective, as the authors argue.

They also acknowledge the missing funding side. The model omits liabilities, deposit outflows and other pressures that brought down Silicon Valley Bank; they present it as one input to asset-liability management. Yet the abstract and conclusion invoke that bank in arguing against static rules. Their policy "will accept additional duration exposure whenever it raises the expected value". Concentrated, unhedged duration was half of what sank that bank. The omitted deposit run was the other half.

Volatility scaling sharpens the concern.

At c = 0.001, the edge over mean-variance changes only from 3.32 to 3.51 points between the low and high regimes. Drawdown climbs from 2.38% to 5.20% across that range. Return barely responds to volatility; risk does. Over twelve months at zero cost, the tail ranking reverses: the policy's fifth percentile of 1.0196 exceeds mean-variance's 1.0032. At c = 0.01, the policy is back below at 0.9864.

Trading through 100bp costs

On the base curve, the dynamic policy still has an edge. Six-month turnover declines from 6.339 at c = 0 to 1.096 at c = 0.01. It still reallocates more than the whole book and beats mean-variance by 0.71 point. Total cost paid reaches its maximum of 1.434% of wealth at c = 0.005, then drops to 1.096% at c = 0.01 as trading volume falls faster than the rate rises.

The starting curve matters more than that base result suggests. From an inverted start, the 100bp edge is 0.05 points (t = 5) with turnover of 0.15. The authors' original start yields 0.10 points with turnover of 0.19. Upward-sloping and near-flat starts retain edges of 0.52 and 0.58 points, on turnover of 1.53 and 1.17.

Treasury trading costs are well under 10bp, making the 100bp case a stress test. At c = 0.0005, the policy turns over 5.815 times in six months and finishes at 1.0829. The more realistic version is near-frictionless and turns its book over roughly six times in half a year.

Scored on the chain used to solve it

The authors say "all results reported here are obtained under the data-generating process assumed by the model itself" and call the figures an upper bound. Section 4.1 goes further in its description of the 10,000 evaluation paths: they are paths of the Markov chain. At minimum, the 40bp grid therefore enters both the optimization and the evaluation. On that reading, the truncation does too. Scoring a policy on its solving chain can preserve an edge created by that chain's artifacts.

Here the paper makes its most useful contribution. With a 20bp grid and 10 retained successors, the chain keeps about a third of the conditional mass and roughly halves each yield's conditional standard deviation. A 40bp grid with 25 successors retains about 92%. The authors also claim truncation leaves mean wealth almost unchanged while distorting fifth-percentile and drawdown statistics. In our reading, the mass and volatility diagnostics support that claim. We did not find a table placing wealth, fifth-percentile and drawdown figures for the two chains side by side. Nor did we find a run scoring the solved policy on continuous VAR paths. Such a run would help separate forecasting skill from grid artifacts.

The wealth levels warrant the same check. The buy-and-hold ladder returns 4.31% in six months, about 8.6% annualized, a large figure for a 2/5/10 Treasury book. Its result varies with the starting curve: 1.0202 from an upward-sloping start and 1.1069 from an inverted one. The inverted start's 10.69% six-month return makes the concern sharper. We could not reconcile these levels because we did not find either the decay parameter λ or the level factor's long-run mean printed in the paper.

We could not run this ourselves. We lack a price series for individual zero-coupon bonds aging along the maturity path, while the paper prices each bond using its shrinking time to maturity. Bond ETFs and Treasury futures maintain roughly constant maturity and roll. Substituting them would test a different strategy.

The authors plan to estimate the model on observed Treasury yields and test it out of sample during the 2021 to 2023 tightening. A multi-point edge over the best single maturity at 5bp costs in that test would make the return claim worth attention. For now, the finding to carry forward is the truncation check: measure retained mass before trusting a binned chain's tail figures.