A posterior-weighted mark can deliver a known loss at a known date on every path, even when every calibrated component is arbitrage-free. Brutsche, Schmidt and Sester construct the point with two dates, two models and one claim. Anyone able to trade against the output earns a guaranteed half a unit.
Why 5/2 becomes 2 on every path
Rates are zero, the bank account remains fixed at 1, and trading takes place at times 0, 1 and 2. The stock begins at S_0 = 2. There are two candidate physical models. The first generates independent Gamma(2,1) prices, while the second uses independent exponentials of rate 1. Both transition densities are strictly positive, leaving every observation compatible with either model.
The claim is the squared one-period return, H = ((S_2 - S_1)/S_1)^2. Each model has its own equivalent martingale measure. Under those measures, log returns are conditionally lognormal with variance sigma_j^2(s). The authors choose sigma_1^2 = log 2 and sigma_2^2(s) = log(3 + s). Since the conditional return variance is exp(sigma_j^2) - 1, the time-1 value of H equals 1 under model 1's measure and 2 + S_1 under model 2's. Their time-0 values are 1 and 4. Starting from the uniform prior in the Dümbgen and Rogers recursion gives a mixed price of 5/2.
The update produces the cancellation. At s, the likelihood ratio between Gamma(2,1) and Exp(1) is exactly s. Model 1 therefore receives posterior weight S_1/(1 + S_1), while model 2 receives 1/(1 + S_1). The mixed conditional value is s/(1+s) times 1, plus 1/(1+s) times (2+s). This reduces to (2s+2)/(1+s). Two, for every realised S_1.
The trade is immediate. Short one unit at 5/2 at time 0, leave the proceeds in the bank account, then repurchase it for 2 at time 1. The trader retains 1/2 from zero initial capital, keeps non-negative wealth at every date, holds no underlying position and needs no view about the correct physical model.
No market test fits the claim
This trade depends on being able to transact at the model's posterior-weighted valuation. No venue quotes such an output. Our data consist of end-of-day listed option prices, which provide nowhere to execute the short-and-cover leg. US equity options are the nearest listed universe available to us. A run on that market would examine a multi-model pricing and calibration setup of our choosing, rather than reproduce the paper's constructed arbitrage.
The authors say the counterexample was constructed with the aid of ChatGPT and report no independent verification. Its algebra takes three lines. It checks out.
Averaging causes the break
The paper gives the right diagnosis in one sentence. Physical likelihoods update the weights, while model-specific risk-neutral measures supply the conditional values. The procedure multiplies two different objects. As the authors observe, taking expectations under the single fixed mixture measure Q1/2 + Q2/2 would preserve dynamic consistency and exclude arbitrage. Its conditional model weights generally depart from the physical posterior weights above.
Each component works on its own. Every model is arbitrage-free and carries an equivalent martingale measure, while all four measures have strictly positive joint densities on (0, infinity)^2.
The abstract states the condition directly: "If these prices are tradable", shorting and covering produces a certain profit. The introduction draws the same boundary. For non-tradable valuation estimates, the example establishes a failure of dynamic pricing consistency. That distinction determines the desk interpretation. When the engine supplies marks, the predictable move from 5/2 to 2 creates a hedge-ratio and risk-reporting problem rather than a hole in the P&L.
How engineered is the result?
Quite a lot.
The choice sigma_2^2(s) = log(3 + s) gives model 2 a conditional value of 2 + S_1. That expression is precisely what the reciprocal posterior weight requires to remove the state dependence and leave a constant. The authors selected this pair of martingale measures. The construction is legitimate because a counterexample needs only one instance. Without exact cancellation, the time-1 mixed price becomes random, and an arbitrage requires it to remain below 5/2 on every path. The example proves that the rule lacks a no-arbitrage guarantee. It leaves the prevalence of such gaps open.
Another assumption is explicit in the paper. The two models are taken to have identical calibration losses, causing the quadratic price-fit term to cancel from the weights. The footnote acknowledges that, under the martingale measures selected in (7), this "typically will only hold for selected derivatives". Setting the loss term to zero makes the equality automatic. The Dümbgen and Rogers score then becomes a pure Bayesian posterior over the underlying transitions. In its full form, the recursion also penalises differences between model and market derivative prices. The paper does not resolve whether a live price-fit term could prevent the deterministic 1/2 gap.
The appendix goes further
The appendix gives a sharper finite-state example (u = 1.1, m = 1, d = 0.9, S_0 = 100, zero rates). In case (ii), equal priors make the posterior weight process a martingale under both risk-neutral measures, with expectation 1/2 under each. Recursive valuation still fails. For the call (S_2 - 100)^+, direct time-zero valuation gives 553/128, whereas valuing the time-one price again at time zero gives 2245/512. The difference is 33/512.
Case (i) starts with priors 1/4 and 3/4. Here the posterior is a martingale under neither measure: its expectation is 11/40 under Q_1 and 23/80 under Q_2, compared with a prior of 1/4. The digital on {S_2 = 121} has a direct value of 31/256 and an intermediate-value result of 55/512. The gap is 7/512. Martingality of the model weights under each pricing measure is therefore insufficient for recursive consistency. That conclusion travels further than the engineered continuous-state decline.
The authors give two caveats, both of which bear repeating. Their trinomial violates the Lebesgue-density assumption in the original framework. Converting its inconsistency into an arbitrage also requires the time-one value process to be a traded claim. The continuous-state construction requires only the one derivative and cash.
A practical failure would become persuasive if the deterministic decline survived a canonical choice of pricing measures, minimal entropy or the like. A live price-fit term in the likelihood would also need to survive. For now, the result is narrower and useful: a posterior-weighted average of arbitrage-free model prices remains a valuation estimate and should never be quoted as a price.