The worst-case correction in this paper does not depend on the preference you were unsure about. Stare at that before anything else. Under a Wasserstein ball around your reference weight function, the worst case is the reference value plus sqrt(eps) times the L2 norm of the loss, and the reference weight function has vanished from the correction term. Two desks holding the same book, one running Expected Shortfall and one running a Wang transform, pay an identical ambiguity charge at the same radius.
The object under study
A generalised distortion risk measure is a weighted average of quantiles: rho_gamma(Y) = int_0^1 F^{-1}_Y(u) gamma(u) du, where the weight function gamma is chosen by the decision maker. Value-at-Risk, Tail Value-at-Risk and the Wang transform are all that object with different gamma. The shape of gamma is where risk preference lives: a steeper gamma weights large quantiles more, which is loss aversion.
Classical robustness perturbs the loss distribution and holds the risk functional fixed. Bernard and Pesenti perturb the weight function instead and hold the distribution fixed. The decision maker names a reference weight function gamma_0, their best guess, and an ambiguity radius eps; the best and worst case are then taken over all weight functions within eps of gamma_0. When revealed preferences are supplied, the authors also characterise eps_min, the smallest radius consistent with those choices. The second-moment charge that runs through the rest of this piece comes out of that perturbation of gamma, not of F_Y.
There is no dataset anywhere in the paper. Every illustration is parametric.
What Bernard and Pesenti actually build
The premise is that gamma itself is uncertain. Survey-based elicitation is inconsistent, and the authors cite that literature directly. Our own reading is that a firm naming an ES level has not thereby committed to a full tail-weighting profile. So the authors take gamma_0 and evaluate the best and worst case over a neighbourhood of it.
Two geometries. Squared L2 distance, which they call Wasserstein because monotone weight functions can be read as quantile functions. And the wider Bregman family, where the generator lets deviations on the left and right of gamma_0 be priced asymmetrically. Both give closed forms. They then add economically motivated restrictions: coherence (gamma non-negative, non-decreasing, integrating to one), revealed pairwise lottery preferences, and an inverse-S shape with an exogenously supplied inflection point (u_dagger = 0.5 in Figure 9). Section 4 applies the same machinery to rank-dependent utility with gamma_0 identically 1, which is expected utility, and shows that RDU emerges as its robustification.
The illustrations are Pareto with alpha = 3 and x_m = 1, a N(0, 2^2), a standard lognormal, constant relative risk aversion utility with r = 2, and the Allais lotteries with payoffs of 1,000,000 and 5,000,000. Figures 2 to 5, which illustrate Theorems 3.1 to 3.5, are drawn at eps = 0.1. Figures 7 and 9, which illustrate the rank-dependent utility results, use eps = 0.01, 0.1 and 1. The paper makes no empirical claim and reports no performance for itself, and the conclusion names estimating and eliciting gamma_0 and eps from data as next steps. We have written before about putting the Wasserstein ball around the distribution rather than around the preference (our note). This is the other side of the same ball.
The paper gives a risk-measure framework and no trading rule, so the reference distortion, the ambiguity radius, the lookback and the position constraints in the numbers printed above this piece are ours. They score a rule we built from Theorem 3.1, not a claim Bernard and Pesenti make.
One formula, and it ignores your reference
Theorem 3.1 is a Cauchy-Schwarz argument, and the result is exact: rho_wc(Y) = rho_gamma0(Y) + sqrt(eps)||Y||_2, rho_bc(Y) = rho_gamma0(Y) minus the same term, with almost everywhere unique extremal weights gamma_0(u) plus or minus (sqrt(eps)/||Y||_2) F^{-1}_Y(u). Summarising this first contribution, the authors put it plainly: preference ambiguity endogenously generates a second moment loading.
For a risk report, this is cheap and honest. One extra statistic, ||Y||_2, and one dial gives you a symmetric band around whatever number you already publish. For anyone ranking positions, the news is thinner. The added term is monotone in ||Y||_2 and carries no tail information the second moment does not already have. Note also that ||Y||_2 is not a dispersion measure, so a book with a large expected loss is charged for its mean as well. Under the assumptions of Theorem 3.3, and for a radius small enough to satisfy condition (3.6), Corollary 3.1 turns the shift into sqrt(eps) times std(Y), a volatility loading bolted onto a tail measure. Corollary 3.1 is the paper's most interpretable result, and the one that should make a tail-risk desk pause: uncertainty about your tail aversion gets priced exactly like variance.
Does the volatility reading survive an ES weight?
Corollary 3.1 holds when sqrt(eps) <= gamma_0(u) std(Y) / (E[Y] - F^{-1}_Y(u)) for every u where the quantile sits below the mean. It also inherits Theorem 3.3's hypothesis that gamma_0 is coherent: non-negative, non-decreasing, integrating to one. The paper's sufficient version of the radius condition asks for gamma_0(u) >= c_1 > 0 across the whole unit interval, Y bounded below, and eps small. An ES weight is zero below the level, gamma_0(u) = (1/(1-alpha)) 1{u > alpha}. The right-hand side of that bound is therefore zero at any u below alpha with F^{-1}_Y(u) < E[Y], and no positive eps clears it. The clean standard-deviation form belongs to smooth, everywhere-positive coherent references such as the power distortion gamma_0(u) = 0.7(1-u)^{-0.3} the paper uses for its coherent illustrations in Figures 4 and 5, chosen, in the authors' words, "to ensure that it is increasing". Run ES and you are in Theorem 3.3's truncated positive-part form instead, where the shift is not a volatility number.
The same shape of problem appears on the utility side, and here the authors flag it themselves: Corollary 4.1's closed form fails for every eps > 0 when U(Y) is unbounded below, which covers CRRA with r >= 1, and their own illustrations use CRRA with r = 2.
Epsilon moves with whatever you anchor on
The calibration proposal is the smallest radius consistent with revealed preferences. The authors concede in their conclusion that estimating and eliciting gamma_0 and eps from data remains future work, and they answer the gap in the same breath by characterising eps_min and calling it a natural choice of radius. At eps_min the ambiguity set collapses to a single weight function, so what comes back is a re-estimated preference rather than a band.
Their financial example makes the tension concrete. W ~ Exp(1) and V ~ N(1, 1.2^2) share a mean of 1, var(W) = 1 against var(V) = 1.44, but ES_0.95(W) is about 4.00 against about 3.48 for V. Mean-variance prefers W, ES prefers V. Start from an ES reference, reveal a preference for W, and the minimal radius is about 0.99, with the squared Wasserstein distance between the two quantile functions at about 0.27. At that radius the set collapses to the L2 projection of the ES weight onto the half-space, and best and worst case coincide. Choose any larger radius and the band reopens, at a width you picked.
The radius also depends on the guess you anchored on. Same two Allais lottery pairs, two references: eps_min is about 0.095 for the proportional-hazard weight with beta = 0.7 and about 0.368 for Prelec with alpha = 0.65, beta = 1, with multipliers around (1.54, 1.71) and (1.82, 2.24) respectively. The RDU version of the same paradox, anchored at gamma_0 identically 1 with CRRA r = 2, gives about 0.063. The calibrated radii reported in the paper run from 0.063 to 0.99. The illustrations of the risk-measure results in Figures 2 to 5 are all drawn at 0.1. Anyone tempted to govern this with a house-wide eps should sit with that gap: the only financial calibration in the paper, eps_min of about 0.99, is roughly ten times the 0.1 used in Figures 2 to 5, while Figures 7, 9 and 10 go as high as eps = 1.
The Allais result holds up. At eps_min of about 0.063, the L2 projection of the expected-utility benchmark onto the Allais-consistent set is, in the paper's phrase, "the closest RDU to expected utility that rationalises the observed choices": it overweights probabilities below 1% and underweights those above 90%. The inverse-S shape itself is not derived from that; Section 4.3 imposes it as a constraint set, with the inflection point supplied by the user (0.5 in Figure 9). The isotonic-projection observation is also worth keeping: with a power-distortion reference with beta = 0.7 and a Pareto-tailed loss, the isotonic projection in (3.9) forces the best-case coherent weight flat on a right tail, so the best-case coherent weight caps tail weighting rather than running to a corner.
Our run, and what it can and cannot say
The paper specifies no trading rule, so the reference distortion, the radius, the lookbacks, the gates and the position caps are all ours. Nothing below tests a claim of theirs, and the paper reports no performance of its own, so the figures printed above stand alone as the output of a rule we specified.
We built a monthly long-only rotation over the top 60 US ETFs by trailing one-year dollar volume, on daily bars, from 2015-01-01 to 2024-12-31. It returned 30.09% in total with a Sharpe of 0.27. Each month we form signed losses Y = -r over 252 trading days (minimum 189 observations), compute rho_gamma0 against the increasing power weight gamma_0(u) = 0.7(1-u)^{-0.3}, the one the paper plots for its coherent cases, and add sqrt(0.1) times the L2 loss norm. Eligible ETFs are then scored as 126-day trailing return divided by that worst-case risk number, with eligibility requiring a positive 126-day return and rho_wc <= 0.03. We hold the top 15 names, weighted inverse to the worst-case risk, capped at 10% each, remainder in cash; portfolio-level gates halve gross above 0.02 and go flat above 0.03. Costs were four tenths of a cent a share with a $1 minimum, and zero modelled slippage.
Three things about our build shape those numbers more than the paper does. We implemented the unconstrained Wasserstein formula, not the coherent or Bregman variants, so nothing here speaks to Theorem 3.3 or Corollary 3.1. Our ranking denominator multiplies the per-day rho_wc by 126, while the sqrt(eps)||Y||_2 term scales like sqrt(126). That is an 11.2x horizon mismatch on the ambiguity component. The multiplier is constant across ETFs, so rankings are unaffected, but the score is not a literal 126-day formula from the paper. And the universe screen uses one-year dollar volume without a point-in-time membership check, so a liquidity-selection tilt is possible.
The gates matter. They mean gross exposure can sit at zero for stretches, which is where the -23.44% max drawdown comes from as much as from the selection rule.
The deeper point is one the paper's own theorem predicts. Since both terms of rho_wc grow with loss magnitude, our ambiguity-loaded denominator behaves close to a downside-volatility scaler, and what our single automated implementation produced is a momentum-over-volatility rotation with vol gating. The win rate was 61.75% against a Sharpe of 0.27, so hit rate is not the weak part. Any faithful implementation of Theorem 3.1 will have that character, because the ambiguity term is a second-moment loading by construction.
Where this earns its keep
As a reporting object, the formula is worth having: sqrt(eps)||Y||_2 is the price of not knowing your own tail aversion, it is exact, and it costs one statistic to compute. As a portfolio tool it needs a radius that someone can defend, and the paper's own evidence is that the radius is not a preference primitive. Elicit eps from a real book's revealed choices, show that it is stable across two plausible reference distortions, and then show the worst-case ranking beating the reference ranking out of sample. That last step is the one that would move me.