A robo-advisor can recover a client's risk aversion from one score only if the client scores prospects the way PreFER requires. Under constant absolute risk aversion (CARA), an error-free score and a too-conservative or too-aggressive flag identify γ. The score must act as a von Neumann acceptance probability. Wang and Wong call it "a theoretical representation of the intensity of client feedback". Their convergence guarantee depends on that reading of the score and on how the client's outlook relates to the advisor's. Neither behavioural assumption is tested on clients.
The authors allow for coarse ratings, sliders and ordinal responses through a noise model. Yet their i.i.d., zero-mean, bounded error is added to a γ already implied by the response. They assume the step that turns a coarse rating into that implied γ. The identification guarantee therefore rests on a mapping they leave unmodelled.
PreFER is an interactive robo-advisor. It recommends an investment prospect, takes a score, revises its estimate of the client's risk aversion and tries again until she accepts. The authors describe the inference of a latent reward from feedback as inverse reinforcement learning. After acceptance, the process carries the learned preference into later periods with new market parameters. This is a theoretical paper: it has no client data, and its illustrations use synthetic binomial parameters and simulated feedback error.
What does the client score?
The advisor announces a market outlook with a binomial stock: up factor u, down factor d, up probability p and cash paying zero. The client scores a proposed investment before its return is realized. That timing limits what prices can tell us about the mechanism. No price series records her judgment of a prospect before its P&L exists.
The underlying utilities are predictable forward performance processes (PFPPs), rolled forward one period at a time without a fixed terminal horizon. The authors randomize the control and apply a Tsallis entropy regularizer with index m = 2 to obtain the PreFER process. Under CARA, forward utility remains exponential, U_t = -C_t e^{-γx}. Its optimal allocation density has a closed form. The density peaks at the classical strategy π0 = log[p(1-q)/((1-p)q)] / (γ(u-d)), where q = (1-d)/(u-d). That peak holds "regardless of the intended exploration level"; lowering the exploration coefficient ρ makes the density taller and sharper around it.
The peak stays put.
For scoring, the authors use a rejection-sampling acceptance ratio: the client's true density at the proposed strategy divided by c times the advisor's guessed density there. Here c is the supremum of that ratio over the exploration interval. They prove the score falls strictly when the guessed γ moves away from the true γ in either direction.
One score, plus a direction
An exact score identifies two candidate values because it is monotone on each side of the true γ. The client's too-conservative or too-aggressive flag selects between them. With an error-free score and a correct flag, one interaction recovers γ. Both signals have to be clean.
To allow imprecision, the authors write the γ implied by each response as γ^e = γ^a + ε. They assume ε is i.i.d., zero mean, has variance σ² and has bounded support.
The averaging guarantee
Proposition 4 compares the Markovian update, which uses only the latest implied γ, with the cumulative moving average (CMA). For the Markovian update, the expected number of trials to acceptance is 1/p_A. The CMA accepts the N-th guess with probability at least 1 - σ²/(N-1) times a sum of two inverse squared distances to the acceptance boundaries. The recommended strategy converges in probability at rate 1/√N. The bound has the shape of Chebyshev's inequality applied to a sample mean of i.i.d. zero-mean errors.
Table 2 shows simulated paths under a shock probability of 0.4. In Seed 2, the Markovian score reaches 99.4% at interaction 4, then falls to 43.1% at interaction 5 and 58.7% at interaction 8. The CMA rises from 87.7% to 99.7% by interaction 9. In Seed 1, the Markovian path falls to 47.8% at interaction 4; from interaction 4 onward, the CMA stays between 93.5% and 97.0%. These paths come from the authors' simulator and display the noise model they assumed. If a client's responses run consistently hot after a drawdown, the error has a nonzero mean. The CMA would then converge at the same 1/√N rate to the wrong γ.
There is also an internal slip. Algorithm 1 prints a geometric mean, (∏γ^e)^{1/n}, for the update. The text and Proposition 4 use an arithmetic average. Those estimators differ.
We could not test the scoring or the updates. Such a test needs client-level scores and conservative/aggressive flags on proposed prospects; we hold neither. Market prices cannot substitute for judgments made before the outcome.
How many interactions will a client tolerate?
An advisor pays in client patience: each new guess takes an interaction, and there is a limit to how many a client will tolerate. In the Monte Carlo study (5×10^6 trials, uniform error, CMA), the guaranteed score falls as the inconsistency bound d_ε rises from 1 to 2. It improves over n = 1, 2, 3 interactions, though each step adds less. The shock experiment raises d_ε from 1 to 1.9. More interactions are needed as shock probability rises and as the threshold moves from 90% to 95%.
The paper gives one synthetic parameterization (u = 1.15, d = 0.9, γ in [3, 7]) for its carry-forward example and does not vary the market across its illustrations.
The shared-outlook assumption
I part with the authors on decay. Their carry-forward step turns an accepted recommendation into an acceptance interval for later periods. The example starts with a client accepting γ = 5 at a 92% threshold. As p2 rises from 0.65 to 0.75, the interval narrows. For a client whose true γ lies at either edge of the carried-forward interval, a guess of 5.5 scores between roughly 0.99 and 0.93, falling as p2 rises.
The authors leave market misspecification for future work. They say it "may affect the implementation of the learned preference without necessarily distorting its identification." Their defence addresses a shared outlook that later proves wrong. By assumption, it excludes a client who privately holds a different p. Because π0 depends on the log-odds term log[p(1-q)/((1-p)q)] divided by γ, such a client will score as though her γ differs. The advisor cannot separate her belief from her risk aversion. Identification is conditioned on a shared outlook, another behavioural assumption.
A design worth testing on clients
The closed-form density earns the paper credit, as does the p2 result: a more favourable market makes the same γ error cost more score. The paper cites menu methods over finite choice sets; scoring a continuous γ one prospect at a time offers a distinct design. Every reported number, though, comes from synthetic parameters and simulated error. A client panel scoring prospects under a stated outlook would change my view if its scores were monotone in the γ gap and its errors centred on zero.