A 0.90 client weight still lets the manager win the trade in this simulated retirement mandate. One manager euro counts as 1.11 client euros. Kurz, Magg, Kollberg, Stricker, Marx, Reinhardt and Dedić make that arithmetic visible, which is the strongest feature of their framework. The weight by itself tells a reviewer little.
Two principals, one instruction
The client delegates to the wealth manager, which in turn delegates to an AI system. The manager serves as the client's agent and the AI's principal. In the authors' model, the AI is a pure executor without utility of its own. What matters is the instruction the manager gives it.
The AI selects a portfolio-workflow pair: the holdings and the service process used to deliver them. Its objective is λ·V_C + (1−λ)·V_W, combining client value with manager value. An admissibility gate first rules out alternatives that fail suitability, conflict rules, mandate limits or required information under FinSA or MiFID II. Each remaining choice must leave both parties at least as well off as a declared reference service. Those minimums are the floors. Concession accounting places ΔC, the client value surrendered relative to the client-best feasible pair, beside the manager's gain. Staff time and fees shape the workflow regardless; recording the manager's interest makes that trade available for review.
Four constructed mandates supply the cases: business-sale liquidity, retirement drawdown, currency liabilities and concentrated holdings. Each has three portfolios and three workflows (standard, advisory, specialist). The reference service holds cash and charges a 0.4% fee. Monthly cash flows include withdrawals, liabilities, currency holdings, zero-coupon bonds and fees. The historical inputs are 144 monthly ECB observations from 2014 to 2025 for Swiss franc/euro (CHF/EUR) rates and AAA euro-area yields. Growth returns, preferences and the three scenario probabilities (0.25, 0.50, 0.25) are authored. Client value is scaled to 10% of initial wealth; manager value, to 1%.
What happens when the instruction goes wrong?
The complete specification reproduces all eight decisions across eight decision states, formed from four mandates with an original and a corrected liability each. Set the omitted liability to zero, and 2 of 8 recommendations breach the original liquidity requirement. Give the manager zero hourly cost and 100 supervisory hours instead of its specified terms, and 2 of 8 breach capacity. Transmit λ = 0.35, and 4 of 8 choices change while every choice remains admissible.
The weight error is harder to catch. A compliance check can identify a liquidity breach. In the business-sale mandate, however, the authorized 0.95 selects growth/standard while 0.35 selects growth/specialist; both clear the original admissibility constraints. Execution error remains zero throughout. The optimiser follows its instruction.
Six of eight states favoured a higher service tier
At the 10:1 scale ratio, ε = 10(1−λ)/λ expresses how many client euros count as one manager euro. At the declared weights, six states selected advisory or specialist service where the client-best pair selected standard:
- Retirement, λ = 0.90, ε = 1.11: the client concedes EUR 1,190; the manager gains EUR 3,062.
- Currency liability, λ = 0.85, ε = 1.76: EUR 2,264 against EUR 10,381.
- Concentrated holdings, λ = 0.98, ε = 0.20: EUR 1,178 against EUR 6,656.
- Business sale, λ = 0.95, ε = 0.53: the selected pair is the client's best, with zero concession.
At 0.98, advisory still wins, giving up EUR 1,178 of client value for EUR 6,656 of manager income. The authors put the point plainly: "A weight close to one can therefore still favour the manager when the manager scale is the finer one." Their proposed review screen would require evidence of client benefit for every positive concession.
The authors acknowledge both the scale choice and the inactive floors. The 10:1 ratio was authored, and it changes the exchange rate alongside λ. Every hard-feasible pair cleared both floors in every simulated state, so the floor mechanism bites only in the toy constructions. A commercially motivated policy could sit in the scales. Reviewers should audit ε first.
The 2022 shock measures the scenario set
The authors hold each selection fixed, then substitute observed 2022 FX and yield paths while retaining their authored growth scenarios. Across eight states, the average client measure falls EUR 77,967 below the cash reference. Manager income exceeds it by EUR 8,816 in every treatment. For one complete case from each family, the forecast and 2022 figures, in EUR thousands and excluding the risk penalty, are business sale +40.2 vs −69.3, retirement +22.5 vs −47.3, currency liability +56.1 vs −89.1, and concentrated holdings +65.6 vs −117.0.
The five-year yield moved from about −0.48% to 2.45%; the stipulated adverse scenario allowed 1.5 percentage points. The authors flag the mismatch. They also note that "No term depends on the market path, so the manager figure would be the same in any year". Rising rates lift the cash reference through its accrual proxy. The exercise therefore judges a fixed fee schedule against a rate shock absent from the scenario set, using a benchmark lifted by that same shock. It grades the scenarios. The authors limit their interpretation to the specified mix of historical and authored paths and argue that review should cover the forecasts and valuations behind the mandate. I accept that lesson. The EUR 77,967 gap tells me more about a +1.5 point adverse case than about the selected services.
We could not replicate it. Valuing the foreign-currency holdings and liabilities requires a CHF/EUR spot path, and pricing the zero-coupon bonds requires ECB AAA yields. We have neither a CHF/EUR nor a euro-yield series to run it on. US-listed ETFs could support a separate test of mandate-aware selection, without the currency liabilities or bond cash flows.
Document reading under test
All three instruction forms produced 32/32 prescribed decisions (28 issuances, 4 holds) and zero issued violations. Plain professional prose matched explicit nested delegation exactly: 20 questions each, costing USD 6.3146 against 6.3162. The explicit joint-objective version asked 26 questions and cost USD 7.3317.
For the narrative cases, the initial interface scored 42, 30 and 22 out of 42 across deterministic, single-agent and multi-agent workflows. The 24-call calculator budget ran out in 8 single-agent and 16 multi-agent runs. With the revised interface, all workflows reached 42/42. The authors say it "was developed from failures on the same cases used for its evaluation," and all 114 revised calculator events returned a single feasible candidate. These runs tested extraction and evidence control. Multi-agent cost 2.64 times single agent; the deterministic workflow incurred zero model charges.
The framework earns its place as bookkeeping for mandate review, without making a portfolio claim. The test that would change my view is the authors' proposed reviewer study: compare the framework with a competent governance checklist on matched defects, including a mistranslated 0.35 weight.