AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Preference ambiguity priced as a second-moment charge you size yourself

Bernard and Pesenti get an exact closed form, but their own calibrated radii run from 0.063 to 0.99

2026-09-08 · 9 min read · US equities and US ETFs using daily price bars.

Reviewing: Preference robust distortion risk measures · Carole Bernard and Silvana M. Pesenti · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Monthly Wasserstein Preference-Robust Downside-Risk ETF Allocation

Backtest period 2015-01-01 to 2024-12-31 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

This is not a replication of the paper (2)

  • The paper provides a theoretical risk-measure framework rather than a fully specified empirical trading rule, so the reference distortion function, ambiguity radius, lookback window, and portfolio constraints must be chosen by the implementer.
  • Any backtest would evaluate a practical allocation rule derived from the framework, not a trading-performance claim made by the paper.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (31)

  • coherent distortion weights: Γcoh:={γ∈Γ|γ≥0, non-decreasing, and ∫_0^1 γ(u)du=1} (invalidates: Coherent Wasserstein and coherent Bregman results do not apply; γwc is not assumed non-negative, non-decreasing, or normalized.)
  • Definition 2.3 Bregman divergence: B_ϕ(z1,z2):=ϕ(z1)−ϕ(z2)−ϕ'(z2)(z1−z2), z1,z2∈R (invalidates: Bregman implementation results do not apply.)
  • Definition 2.4 integrated Bregman divergence: B_ϕ(γ1,γ2):=∫_0^1 B_ϕ(γ1(u),γ2(u))du (invalidates: Bregman implementation results do not apply.)
  • Definition 2.4 Bregman ambiguity set: B_ε(γ0):={γ∈Γ | B_ϕ(γ,γ0)≤ε} (invalidates: Bregman ambiguity-set results do not apply.)

27 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2015Total 30.1%Dec 2024
Sharpe
0.27
Total Return
30.1%
Max Drawdown
-23.4%
CAGR
2.7%
Volatility
11.2%
Trades
2,044

The worst-case correction in this paper does not depend on the preference you were unsure about. Stare at that before anything else. Under a Wasserstein ball around your reference weight function, the worst case is the reference value plus sqrt(eps) times the L2 norm of the loss, and the reference weight function has vanished from the correction term. Two desks holding the same book, one running Expected Shortfall and one running a Wang transform, pay an identical ambiguity charge at the same radius.

The object under study

A generalised distortion risk measure is a weighted average of quantiles: rho_gamma(Y) = int_0^1 F^{-1}_Y(u) gamma(u) du, where the weight function gamma is chosen by the decision maker. Value-at-Risk, Tail Value-at-Risk and the Wang transform are all that object with different gamma. The shape of gamma is where risk preference lives: a steeper gamma weights large quantiles more, which is loss aversion.

Classical robustness perturbs the loss distribution and holds the risk functional fixed. Bernard and Pesenti perturb the weight function instead and hold the distribution fixed. The decision maker names a reference weight function gamma_0, their best guess, and an ambiguity radius eps; the best and worst case are then taken over all weight functions within eps of gamma_0. When revealed preferences are supplied, the authors also characterise eps_min, the smallest radius consistent with those choices. The second-moment charge that runs through the rest of this piece comes out of that perturbation of gamma, not of F_Y.

There is no dataset anywhere in the paper. Every illustration is parametric.

What Bernard and Pesenti actually build

The premise is that gamma itself is uncertain. Survey-based elicitation is inconsistent, and the authors cite that literature directly. Our own reading is that a firm naming an ES level has not thereby committed to a full tail-weighting profile. So the authors take gamma_0 and evaluate the best and worst case over a neighbourhood of it.

Two geometries. Squared L2 distance, which they call Wasserstein because monotone weight functions can be read as quantile functions. And the wider Bregman family, where the generator lets deviations on the left and right of gamma_0 be priced asymmetrically. Both give closed forms. They then add economically motivated restrictions: coherence (gamma non-negative, non-decreasing, integrating to one), revealed pairwise lottery preferences, and an inverse-S shape with an exogenously supplied inflection point (u_dagger = 0.5 in Figure 9). Section 4 applies the same machinery to rank-dependent utility with gamma_0 identically 1, which is expected utility, and shows that RDU emerges as its robustification.

The illustrations are Pareto with alpha = 3 and x_m = 1, a N(0, 2^2), a standard lognormal, constant relative risk aversion utility with r = 2, and the Allais lotteries with payoffs of 1,000,000 and 5,000,000. Figures 2 to 5, which illustrate Theorems 3.1 to 3.5, are drawn at eps = 0.1. Figures 7 and 9, which illustrate the rank-dependent utility results, use eps = 0.01, 0.1 and 1. The paper makes no empirical claim and reports no performance for itself, and the conclusion names estimating and eliciting gamma_0 and eps from data as next steps. We have written before about putting the Wasserstein ball around the distribution rather than around the preference (our note). This is the other side of the same ball.

The paper gives a risk-measure framework and no trading rule, so the reference distortion, the ambiguity radius, the lookback and the position constraints in the numbers printed above this piece are ours. They score a rule we built from Theorem 3.1, not a claim Bernard and Pesenti make.

One formula, and it ignores your reference

Theorem 3.1 is a Cauchy-Schwarz argument, and the result is exact: rho_wc(Y) = rho_gamma0(Y) + sqrt(eps)||Y||_2, rho_bc(Y) = rho_gamma0(Y) minus the same term, with almost everywhere unique extremal weights gamma_0(u) plus or minus (sqrt(eps)/||Y||_2) F^{-1}_Y(u). Summarising this first contribution, the authors put it plainly: preference ambiguity endogenously generates a second moment loading.

For a risk report, this is cheap and honest. One extra statistic, ||Y||_2, and one dial gives you a symmetric band around whatever number you already publish. For anyone ranking positions, the news is thinner. The added term is monotone in ||Y||_2 and carries no tail information the second moment does not already have. Note also that ||Y||_2 is not a dispersion measure, so a book with a large expected loss is charged for its mean as well. Under the assumptions of Theorem 3.3, and for a radius small enough to satisfy condition (3.6), Corollary 3.1 turns the shift into sqrt(eps) times std(Y), a volatility loading bolted onto a tail measure. Corollary 3.1 is the paper's most interpretable result, and the one that should make a tail-risk desk pause: uncertainty about your tail aversion gets priced exactly like variance.

Does the volatility reading survive an ES weight?

Corollary 3.1 holds when sqrt(eps) <= gamma_0(u) std(Y) / (E[Y] - F^{-1}_Y(u)) for every u where the quantile sits below the mean. It also inherits Theorem 3.3's hypothesis that gamma_0 is coherent: non-negative, non-decreasing, integrating to one. The paper's sufficient version of the radius condition asks for gamma_0(u) >= c_1 > 0 across the whole unit interval, Y bounded below, and eps small. An ES weight is zero below the level, gamma_0(u) = (1/(1-alpha)) 1{u > alpha}. The right-hand side of that bound is therefore zero at any u below alpha with F^{-1}_Y(u) < E[Y], and no positive eps clears it. The clean standard-deviation form belongs to smooth, everywhere-positive coherent references such as the power distortion gamma_0(u) = 0.7(1-u)^{-0.3} the paper uses for its coherent illustrations in Figures 4 and 5, chosen, in the authors' words, "to ensure that it is increasing". Run ES and you are in Theorem 3.3's truncated positive-part form instead, where the shift is not a volatility number.

The same shape of problem appears on the utility side, and here the authors flag it themselves: Corollary 4.1's closed form fails for every eps > 0 when U(Y) is unbounded below, which covers CRRA with r >= 1, and their own illustrations use CRRA with r = 2.

Epsilon moves with whatever you anchor on

The calibration proposal is the smallest radius consistent with revealed preferences. The authors concede in their conclusion that estimating and eliciting gamma_0 and eps from data remains future work, and they answer the gap in the same breath by characterising eps_min and calling it a natural choice of radius. At eps_min the ambiguity set collapses to a single weight function, so what comes back is a re-estimated preference rather than a band.

Their financial example makes the tension concrete. W ~ Exp(1) and V ~ N(1, 1.2^2) share a mean of 1, var(W) = 1 against var(V) = 1.44, but ES_0.95(W) is about 4.00 against about 3.48 for V. Mean-variance prefers W, ES prefers V. Start from an ES reference, reveal a preference for W, and the minimal radius is about 0.99, with the squared Wasserstein distance between the two quantile functions at about 0.27. At that radius the set collapses to the L2 projection of the ES weight onto the half-space, and best and worst case coincide. Choose any larger radius and the band reopens, at a width you picked.

The radius also depends on the guess you anchored on. Same two Allais lottery pairs, two references: eps_min is about 0.095 for the proportional-hazard weight with beta = 0.7 and about 0.368 for Prelec with alpha = 0.65, beta = 1, with multipliers around (1.54, 1.71) and (1.82, 2.24) respectively. The RDU version of the same paradox, anchored at gamma_0 identically 1 with CRRA r = 2, gives about 0.063. The calibrated radii reported in the paper run from 0.063 to 0.99. The illustrations of the risk-measure results in Figures 2 to 5 are all drawn at 0.1. Anyone tempted to govern this with a house-wide eps should sit with that gap: the only financial calibration in the paper, eps_min of about 0.99, is roughly ten times the 0.1 used in Figures 2 to 5, while Figures 7, 9 and 10 go as high as eps = 1.

The Allais result holds up. At eps_min of about 0.063, the L2 projection of the expected-utility benchmark onto the Allais-consistent set is, in the paper's phrase, "the closest RDU to expected utility that rationalises the observed choices": it overweights probabilities below 1% and underweights those above 90%. The inverse-S shape itself is not derived from that; Section 4.3 imposes it as a constraint set, with the inflection point supplied by the user (0.5 in Figure 9). The isotonic-projection observation is also worth keeping: with a power-distortion reference with beta = 0.7 and a Pareto-tailed loss, the isotonic projection in (3.9) forces the best-case coherent weight flat on a right tail, so the best-case coherent weight caps tail weighting rather than running to a corner.

Our run, and what it can and cannot say

The paper specifies no trading rule, so the reference distortion, the radius, the lookbacks, the gates and the position caps are all ours. Nothing below tests a claim of theirs, and the paper reports no performance of its own, so the figures printed above stand alone as the output of a rule we specified.

We built a monthly long-only rotation over the top 60 US ETFs by trailing one-year dollar volume, on daily bars, from 2015-01-01 to 2024-12-31. It returned 30.09% in total with a Sharpe of 0.27. Each month we form signed losses Y = -r over 252 trading days (minimum 189 observations), compute rho_gamma0 against the increasing power weight gamma_0(u) = 0.7(1-u)^{-0.3}, the one the paper plots for its coherent cases, and add sqrt(0.1) times the L2 loss norm. Eligible ETFs are then scored as 126-day trailing return divided by that worst-case risk number, with eligibility requiring a positive 126-day return and rho_wc <= 0.03. We hold the top 15 names, weighted inverse to the worst-case risk, capped at 10% each, remainder in cash; portfolio-level gates halve gross above 0.02 and go flat above 0.03. Costs were four tenths of a cent a share with a $1 minimum, and zero modelled slippage.

Three things about our build shape those numbers more than the paper does. We implemented the unconstrained Wasserstein formula, not the coherent or Bregman variants, so nothing here speaks to Theorem 3.3 or Corollary 3.1. Our ranking denominator multiplies the per-day rho_wc by 126, while the sqrt(eps)||Y||_2 term scales like sqrt(126). That is an 11.2x horizon mismatch on the ambiguity component. The multiplier is constant across ETFs, so rankings are unaffected, but the score is not a literal 126-day formula from the paper. And the universe screen uses one-year dollar volume without a point-in-time membership check, so a liquidity-selection tilt is possible.

The gates matter. They mean gross exposure can sit at zero for stretches, which is where the -23.44% max drawdown comes from as much as from the selection rule.

The deeper point is one the paper's own theorem predicts. Since both terms of rho_wc grow with loss magnitude, our ambiguity-loaded denominator behaves close to a downside-volatility scaler, and what our single automated implementation produced is a momentum-over-volatility rotation with vol gating. The win rate was 61.75% against a Sharpe of 0.27, so hit rate is not the weak part. Any faithful implementation of Theorem 3.1 will have that character, because the ambiguity term is a second-moment loading by construction.

Where this earns its keep

As a reporting object, the formula is worth having: sqrt(eps)||Y||_2 is the price of not knowing your own tail aversion, it is exact, and it costs one statistic to compute. As a portfolio tool it needs a radius that someone can defend, and the paper's own evidence is that the radius is not a preference primitive. Elicit eps from a real book's revealed choices, show that it is stable across two plausible reference distortions, and then show the worst-case ranking beating the reference ranking out of sample. That last step is the one that would move me.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each last trading day of month t:
  For each ETF in the top-60 dollar-volume universe:
    Build close-to-close daily returns over the trailing 252 trading days.
    Require at least 189 observed returns.
    Define signed loss Y_s = -r_s, with gains entering as negative losses.
    Estimate empirical quantiles of Y and reference distortion weight gamma0(u)=0.7*(1-u)^(-0.3).
    Compute reference distortion risk rho_gamma0(Y).
    Compute L2 loss norm = sqrt(mean(Y_s^2)).
    Compute Wasserstein worst-case risk rho_wc = rho_gamma0 + sqrt(0.1)*L2_norm.
    Compute trailing_return = close_t / close_{t-126} - 1.
    ETF is eligible only if trailing_return &gt; 0 and rho_wc &lt;= 0.03.
    Score = trailing_return / max(126 * rho_wc, 0.000001).

  Rank eligible ETFs by descending score and keep the top 15.
  Set raw_weight_i = 1 / max(rho_wc_i, 0.000001).
  Normalize to target gross exposure of 1.0, then cap each ETF at 10%; unallocated capital stays in cash.
  Using proposed component weights and historical component returns, estimate portfolio rho_wc.
  If portfolio rho_wc &gt; 0.03: target gross exposure = 0 and hold cash.
  Else if portfolio rho_wc &gt; 0.02: target gross exposure = 0.5 and rescale weights.
  Execute required trades at the same rebalance close using observed close prices; skip orders with missing close data.