230 is the figure that matters. Noguer i Alonso reports a 59.64 times increase in certainty equivalent after quadratic path functionals enter a cross section of twenty assets. Yet his word-class decomposition assigns the entire level-two gain to convexity in the terminal log return. Genuinely path-dependent words add exactly nothing. The useful result for a desk lies underneath: 230 generator parameters recover 0.950 of oracle value with barely one observation per parameter, while freely estimating the risk form produces minus 306 times that value.
The paper makes no empirical claim. The author says so directly: "No market data is used anywhere. All parameters are invented and internally consistent." Everything that follows is theory and Monte Carlo applied to drivers chosen by the author.
The framework and its structural results
The portfolio is a linear functional of the truncated path signature. Its control therefore becomes a coefficient tensor instead of a weight vector. Level one reproduces buy and hold. At level two, the words are iterated integrals of one asset against another. Matching indices encode momentum; different indices encode lead-lag.
Risk enters through the defect form, defined as D(u,v) = ⟨u ⊔⊔ v, E[S]⟩ minus ⟨u,E[S]⟩⟨v,E[S]⟩. This is the non-group-like part of the expected signature. Noguer i Alonso proves that it equals the covariance of signature coordinates. Mean-variance then reduces to a single linear solve, ℓ* = D⁻¹µ/γ, and the certainty equivalent is µ'D⁻¹µ/(2γ). When E[S] is known, the wider basis pays. Oracle certainty equivalent rises from 0.015008 at level one to 0.167342 for d=2, and from 0.027984 to 1.669016 for d=20.
Two identities would remain useful even if every numerical result disappeared. At level two, the difference between the Marcus and forward lifts is one half the quadratic covariation. Contracting that difference with diag(π) minus ππ' gives Fernholz's excess growth rate exactly. The rebalancing premium is therefore the lift gap, fixed by the assumed execution convention rather than the dynamics alone.
The second identity matters directly to a risk system. A jump cannot extinguish Marcus-lifted wealth, while forward-lifted wealth can be extinguished. The forward lift adds d(d+1)/2 directions, precisely the quadratic-covariation payoffs. Variance and covariance swaps consequently fall outside the geometric-lift investable universe. At d=4, with N₂=20 and L₂=10, both identities hold to 2.2×10⁻¹⁵.
The estimation study uses two calibrations of X_t = bt + σW_t on [0,1], under a geometric lift with γ=3. The pair has p=6 and cond D = 134.3. For the cross section, volatilities are equally spaced on [0.22,0.35], equicorrelation is 0.40, drifts span [0.04,0.10], p=420 and cond D = 1906. From M independent paths, he estimates µ̂ and D̂, constructs the plug-in ℓ̂, then evaluates its true certainty equivalent against the exact population (µ,D). Separate bivariate Hawkes simulations handle the lead-lag trade.
Where does the sixtyfold gain come from?
Adding the antisymmetric level-two words, the Lévy-area coordinates, changes certainty equivalent by 1.00 times relative to level one. The result holds for d=2 and d=20, with agreement to sixteen digits. Every gain comes from the symmetric block.
For d=2, the diagonal words account for 10.71 of the 11.15 times. At d=20, the diagonal block contributes 14.06 and the symmetric off-diagonal block 2.91; together they produce 59.64. Under the geometric lift, the shuffle relation makes this symmetric block exactly ½ΔX⊗ΔX. These words are quadratic functions of the terminal log return. In the d=2 solution, the optimizer assigns 7.9375 to the (11) word and numerically identical weights of −1.1689 to (12) and (21). Its optimal control therefore has zero Lévy-area component.
Noguer i Alonso states the source plainly. The paper says that "the honest name for the 59.64× is therefore the price of admitting quadratic payoffs in the terminal increment," and the abstract concedes the same point alongside the headline. Contribution 3 describes the premium as dimensional and locates it in the symmetric block. The paper also observes that quadratic directions increase from three to two hundred and ten while value rises tenfold. With equicorrelated assets, each additional quadratic direction is worth little, while estimation cost grows with their number.
The author offers a defence: "Section 8 is where the antisymmetric words are put in front of a driver that pays them." He also gives the zero a narrow interpretation. In his words, the antisymmetric block earning exactly zero "is a statement about the driver, which is time-reversible and has no expected area, and not a statement about markets."
My objection survives those concessions. The two components are priced under separate drivers. Section 8 "measures that separately rather than jointly", as the paper puts it, and sets the sign-copying probability to q=0.85 instead of estimating it. The title names a dimensional estimation cost; the headline multiple comes from convexity in the terminal increment. The paper acknowledges both facts, though in different places.
Minus 306 times oracle
At d=20, p=420 and M/p=1.19, the median realized certainty equivalent from the unregularized plug-in is −306.1 times its target oracle. The interquartile range is [−634,−402] in absolute terms, with Pr[CE<0] = 1.00. At M/p=2.38, the median remains negative at −2.311. It first turns positive at 4.76, reaching 0.405 of oracle. The plug-in crosses half the oracle value between M/p of 4.8 and 7.1, then reaches 0.973 at M/p=47.6. Across that same range, the level-one plug-in returns 0.690 of its own level-one oracle at M=500 and 0.993 at M/p=47.6.
Until enough data arrive, the wider basis destroys value.
Width also reverses the effect of shrinkage. With fixed ridge at δ=0.25 and p=420, performance ranges from 0.577 to 0.629 of oracle. At M=500, that changes −306 times into plus 0.58. For p=6, the same intensity produces 0.485 to 0.503, versus 0.998 from the raw plug-in at M=8000. Fixed ridge never converges and stays flat across every row in both tables. The Marchenko-Pastur radius δ=√(p/M) starts worse at the smallest sample, 0.301 against 0.577, yet it is the only ridge that converges. It rises to 0.756 at p=420 and 0.895 at p=6.
Noguer i Alonso gives two caveats. Neither intensity was tuned, so the reported fractions fall short of the best achievable. The Marchenko-Pastur localization is also a located analogy instead of a theorem about D̂, since signature coordinates are neither independent nor identically distributed across levels. His separate use of Kan-Zhou exact finite-sample theory applies exactly only to the level-one sub-problem. Level-two payoffs are quadratic in Gaussians, and D̂ is not Wishart. The negative values in Table 2 are therefore measured rather than derived.
The paper is equally direct about the distance between the theorem and the tables. Its theorem assumes a good event in which the sample covariance resolves the spectrum. A sample of size M≈p does not supply that event. In the tables, λmin(D̂) fails, while λmin(D) does not. The theorem bounds losses after the sample resolves the spectrum; the tables record the earlier regime.
Replication counts decline from 200 to 20 across the d=20 rows. The intervals that look tightest rely on the fewest draws.
Structure repairs the estimate
The model-consistent estimator fits the driver generator (b̂,Σ̂) from level-one increments, then reconstructs µ̂ and D̂ through tensor exponential and shuffle. At d=20, this requires 230 numbers. Free estimation uses 420 means and 88,410 covariance entries. The structured version recovers 0.950 of oracle at M/p=1.19 and 0.999 at 47.6. At d=2 with M=60, it reaches 0.957. Against the −306 result, this is the paper's one finding I would act on.
Its scope is also stated by the author. The model-consistent estimator is tested under a correctly specified driver, measuring the cost of declining to use a model while leaving the cost of a wrong model unmeasured. No bound on its bias under misspecification. In an actual market, the generator is the unknown object.
Our adaptation and its limits
We could not reproduce the paper's experiment. Oracle certainty equivalents require E[S] to be known, while model-consistent recovery requires a generator that is true by assumption. Prices reveal neither object.
Because the paper uses generic simulated assets, we substituted a real universe: twenty liquid US stocks and ETFs, using daily closes from 2021-01-01 to 2025-10-08. Our traded basis combines level one with the unique symmetric level-two words. It has p=230 coefficients, a different 230 from the paper's generator count. We set risk aversion γ=3 and use a 1260-day rolling window. Weights are capped at 10% per name with gross 1.0. Orders execute market-on-close, with commissions of $0.004 a share and a $1 order minimum. This exercise adapts the idea and cannot serve as a test of the paper.
Over that period, our run returned 21.36% on 9.96% volatility, with a Sharpe of 0.45. Maximum drawdown was −19.46% across 4542 trades. Nearly twenty percent of drawdown for twenty-one percent of cumulative return leaves a thin result.
The paper reports only simulated, gross Sharpe-like figures. Its sign-carrying cross-area trade earns +1.27 annualized with q=0.85 on a bivariate Hawkes driver; mirrored excitation earns −1.50. Our 0.45 measures a different object from its 1.27. The paper's 1.27 comes from antisymmetric words, which our construction excludes entirely. On a zero-area driver, the paper finds those words earn nothing.
Using non-overlapping windows gives us about 1.1 observations per parameter weekly and 0.26 monthly. Those ratios fall at or below the point where the paper's unregularized plug-in returns −306.1 times oracle for d=20 and p=420. The paper charges no costs, though it identifies the level-two words as exactly the high-turnover ones. We encountered the same cost-threshold problem in our note on fast trend.
Our universe includes QQQ, TQQQ and SQQQ alongside SPY, IVV and VOO. Its defect form is therefore very likely much worse conditioned than the paper's cross-sectional calibration, which has cond D = 1906 at p=420, 81.0% of eigenvalues above 10⁻³λmax, and "the population problem is well posed at both sizes". Any ridge strong enough to resolve that collinearity also dilutes the convexity signal. The weak result speaks only to our construction.
A warning that travels
The cross-area experiments contain one result worth carrying beyond this paper. Excitation that increases intensity without carrying direction shifts the counting path's cross-area. Across 3000 paths at T=20, the mean is +3.253, s.e. 0.218 and t=14.94. The price path shows no corresponding movement. There t=0.96 in the traded version, using tick 0.01 and 2000 paths per row. The paper's contribution list reports t=15.99 for the counting-path leg; 15.99 is Table 5's sign-carrying price-path figure at q=0.85. Reflexivity in activity differs from reflexivity in direction. Only directional reflexivity is tradable.
A related experiment shows the cost of substituting the stationary Hawkes limit. At branching ratio 0.9, it overstates variance by 4.9 times, giving 10000.00 against a simulated 2048.47. Integrating the same closure comes within 1.6% over a horizon of two relaxation times.
One experiment would change my view of the headline. Give the optimizer a single driver containing both expected area and quadratic structure. Estimate q instead of imposing 0.85. Then run the word-class attribution in that shared environment. If the antisymmetric block captures a fraction of the 59.64 there, path complexity has a real price. Until then, the number prices a quadratic payoff basis, exactly as the author has already written.