AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Drawdown-targeted Kelly overlays depend on unobservable drift

Bermin and Holm derive clean geometry, while their 0.665 Sharpe remains a full-sample plug-in

2026-09-18 · 7 min read · US equities and US ETFs

Reviewing: Leverage, drawdowns and risk relativity · Hans‐Peter Bermin and Magnus Holm · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

SPY-Relative Drawdown-Targeted Constrained Kelly Overlay

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

The paper reports no results of its own

This is a theoretical paper — derivations and proofs, with no measurement on market data. The backtest below is a strategy we built from its idea, not a test of anything the authors claimed.

This is not a replication of the paper (4)

  • The theoretical framework assumes continuous trading, frictionless self-financing portfolios, unrestricted short-selling, and no market impact; the platform can only test discretely rebalanced implementations using daily or minute bars.
  • Instantaneous expected returns and covariance processes are unobservable from historical OHLCV data and must be replaced by rolling or exponentially weighted estimates.
  • Kelly-style weights can be unstable and highly leveraged in finite samples; practical implementation requires leverage, borrow, concentration, liquidity, and turnover constraints that are not part of the theoretical model.
  • No equity bid-ask or borrowing-fee data are available, so transaction costs and shorting costs must be modeled rather than directly measured.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (29)

  • Lemma 2 nonzero-aversion lower-barrier probability (invalidates: Exact lower-barrier probabilities are not guaranteed in the daily constrained backtest.)
  • Lemma 2 zero-aversion continuous limit (invalidates: The zero-aversion continuous-path hitting probability is not guaranteed in the daily constrained backtest.)
  • Corollary 3 infinite-horizon drawdown law (invalidates: The exponential infinite-horizon drawdown law and its expectations are theoretical benchmarks rather than exact backtest identities.)
  • Corollary 3 maximum relative loss law (invalidates: The exact maximum-relative-loss survival function and limiting terminal-drawdown quantile are not guaranteed.)

25 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 59.2%Jul 2024
Sharpe
0.67
Total Return
59.2%
Max Drawdown
-35.9%
CAGR
10.9%
Volatility
19.8%
Beta vs SPY
0.83
Trades
1,872

What the paper reports for its own strategy

  • μ_{0|w} = 0.145 with σ_{0|w} = 0.274 and s_{0|w} = 0.665 for the scaled growth-optimal Kelly strategy k_{0|e1}w* on the SPY/GLD/bank opportunity set, computed from constant parameters estimated Nov 2004 – Nov 2020; no transaction costs assumed (paper assumes zero fees, no market impact)
  • μ_{e1|w} = 0.096 with σ_{e1|w} = 0.223 and s_{e1|w} = 0.543 for w = e1 + k_{0|e1}w*_{e1} (index-relative outperformance), same estimates, no costs
  • μ_{e1|w} = 0.032 with σ_{e1|w} = 0.063, s_{e1|w} = 0.543 and R_{e1|w} = 0.061 for w = e1 + k w*_{e1}, k = 0.116, same estimates, no costs
  • Benchmark for comparison within the same table: S&P 500 alone μ_{0|e1} = 0.074, σ_{0|e1} = 0.196, s_{0|e1} = 0.476

A drawdown budget can size an active book, but Bermin and Holm's allocation rests on a number no trader can observe. Their closed form is exact. For any positive benchmark-relative drawdown target, the active position consists entirely of the growth-optimal Kelly portfolio, with the remainder held in the benchmark.

The paper begins with a portfolio measured against a reference asset. Relative drawdown aversion is A = 2μ/σ², where μ denotes the instantaneous log drift of the wealth ratio and σ² its variance rate. Its reciprocal gives the relative drawdown-risk index, R = 1/A. Cash measured against cash has R = 0.

Proposition 4 covers positions with a strictly positive instantaneous rate of return. Across multipliers from 0 to A + 1, R rises strictly with leverage while the position remains drawdown-averse. Beyond k = A + 1, the index changes sign. The authors interpret that region as one where upside barrier events have greater probability than downside barrier events.

The growth-optimal Kelly portfolio is w* = V⁻¹b. Here V is the instantaneous covariance matrix and b is the vector of instantaneous rates of return. This portfolio maximises the instantaneous rate of logarithmic return. Against every reference strategy except itself, it has R = 1, and no other strategy shares that invariance. Measure another portfolio from w* and R = −1 throughout.

From this setup comes the drawdown law. With constant positive A, the infinite-horizon maximum relative loss from initial wealth follows P(L > y) = (1 − y)^(1/R). A single dimensionless value therefore determines the full distribution for a strategy whose positive relative drawdown aversion stays constant. Expected minimum wealth ratio over an infinite horizon equals A/(A+1), while expected log drawdown equals 1/A.

A target R maps to a Kelly loading k = 2R/(1+R). The resulting position is w = u + k(w* − u), with u as the benchmark. R = 1 gives full Kelly relative to that benchmark. For two Kelly strategies measured from a shared reference u, composition satisfies R(w1,w2) = (R(u,w2) − R(u,w1))/(1 − R(u,w2)R(u,w1)) on the non-singular domain. The algebra matches one-dimensional velocity addition, with w* occupying the invariant position. The authors describe the analogy as geometric only.

No backtest appears in the paper, nor does it report realised return or Sharpe figures. Fig. 1 is its sole contact with realised data. It charts the 5 and 10 day drawdown distributions for Kelly multipliers of 1/4, 1/2 and 3/4. The authors call it "an empirical illustration rather than a formal specification test."

The plotted values are analytic local characteristics based on constant parameters estimated across the full Nov 2004 to Nov 2020 sample. The assets are SPY, GLD and a dollar bank account. For the S&P 500, the parameters are μ = 0.0741 and σ = 0.196. Gold has μ = 0.0717 and σ = 0.183, with correlation 0.0394. The theory should be judged on those terms.

Where 0.665 comes from

Using the authors' estimates, the S&P 500 arithmetic rate is 0.0741 + 0.196²/2 = 0.0933. Its Sharpe against the bank account is therefore 0.476, matching the paper. Gold produces 0.483 by our arithmetic; the paper reports no gold Sharpe.

Nearly identical Sharpes and correlation of 0.0394 yield a maximal Sharpe of 0.665, reported as s(0|w*). The increase from 0.476 to 0.665 is diversification arithmetic for a near-orthogonal pair. Its inputs come from the same sixteen years used to calculate the result.

Those conditions leave the Kelly mix especially sensitive to estimation. Gold volatility is 0.183, and 16 years of observations put the standard error of its annualised drift near 4.6 points. The estimate itself is 7.2. For the S&P 500, our arithmetic gives a standard error of roughly 4.9 against 7.41. Relative weights in V⁻¹b are driven by the gap between drifts of 7.2 plus or minus 4.6 and 7.41 plus or minus 4.9.

The framework does not address this problem. Its conclusion lists continuous Brownian paths, frictionless continuous trading, unrestricted short selling, zero costs and market impact, and pointwise instantaneous optimization among the paper's limitations. Estimation error is missing from that list. We raised a related issue when reviewing a Monte-Carlo reconstruction of drawdown tables (note). It matters more here because μ enters directly into the definition of R.

Table 1 makes the exchange plain. Measured against the bank account, the S&P 500 alone has R = 0.259, log return 0.074 and volatility 0.196. Scale w* to the same R and log return reaches 0.145, with volatility 0.274 and Sharpe 0.665.

The overlay w = e1 + k·w*(e1) uses k = 0.116 to match the index's R of 0.259. It records log return 0.106, volatility 0.235 and Sharpe 0.570. Its benchmark-relative index falls to 0.061, accompanied by a benchmark-relative Sharpe of 0.543. Matching the index's absolute drawdown index and exceeding its log return therefore adds 3.9 points of volatility. The geometry holds. The numerical result remains a plug-in estimate.

What does the barrier law add?

The Pareto drawdown law requires constant A. Keeping 2μ/σ² fixed entails continuous exposure adjustments whenever estimated drift or variance changes. The same estimation difficulty returns, now with a feedback loop. Equation 15, which supplies the finite-horizon version plotted in Fig. 1, requires μ and σ themselves to remain constant.

The infinite-horizon identity is the clean result.

Its formulas also exclude or handle by limits the singular cases b = 0, A = 0, k = 2 and R = −1. High leverage approaches those boundaries. The paper prefers a losing Kelly strategy at k = 1 + δ, where δ > 1, to one at k = 1 − δ. Both have identical negative log drift, while the higher volatility raises the probability of reaching the upper barrier first.

The actual comparison sets a short position against one levered beyond 2x. A maintenance call would already have closed the levered book. Bankruptcy is the model's sole absorbing state; a levered ETF position has others. The authors explicitly say R is not a monetary risk measure. They also limit its convexity to strategies that share a Kelly multiplier, constraining its use for aggregation.

Our discrete construction

Direct testing runs into the model's assumptions: continuous trading, frictionless self-financing, unrestricted shorting and no market impact. Daily bars permit only discrete rebalancing. Rolling estimates must replace unobservable instantaneous drift and covariance. Finite-sample unconstrained Kelly weights are unstable, forcing caps that the model does not contain. We lack equity bid-ask and borrow data, leaving shorting and financing costs unmeasured.

Our construction applies the paper's rule to SPY and GLD, plus implicit zero-return cash, each day from 2020-01-01 to 2024-07-01. Drift and covariance use a 252-session rolling window, annualised at 252. The covariance receives 20% diagonal shrinkage and a small jitter term.

We solve for w* across a gross-leverage scenario grid of 1.0, 2.0 and 4.0. Per-asset bounds are ±4, with a 25% ex-ante volatility cap. We then form the overlay w = e1 + k(w* − e1) at k values of 1/4, 1/2 and 3/4. These map to R targets of 1/7, 1/3 and 3/5.

Realised maximum drawdown was -35.92%. This is the outcome of our construction, rather than a test of the (1 − y)^(1/R) law. We skip a rebalance whenever any active weight exceeds 0.50 absolute, projected turnover exceeds 1.0, or the volatility or leverage constraint becomes infeasible. Fills occur at the closing auction. Every fill incurs commissions of $0.0040 a share and a $1.00 order minimum. Borrow, margin financing and market impact are not charged.

The headline figures accompanying this piece come from one automated pass. Both the estimator and window are ours. A 252-day rolling estimate from daily bars over 2020 to 2024 differs from the paper's constant full-sample parameters covering Nov 2004 to Nov 2020. The paper publishes no realised return or Sharpe figures.

Its scaled-Kelly row, k(0|e1)w* at R = 0.259, gives log return 0.145, volatility 0.274 and Sharpe 0.665 with zero fees. Our run delivered annualised volatility of 19.85%, net of commissions, from 2020-01-01 to 2024-07-01. Their value is an instantaneous characteristic calculated from constant full-sample parameters. Ours is realised over four and a half years. The measurements are different objects.

Any movement in our curve speaks first to our implementation choices. The 252-day drift window crosses both the 2020 crash and 2022. When estimates jump, the active cap and turnover gate keep the overlay out of the market. A benchmark-anchored construction using e1 = SPY produced a beta of 0.83 to SPY, direct evidence of that gating.

One experiment would change our view. R-targeting, volatility targeting and fixed fractional Kelly should be compared side by side, each estimated out of sample with rolling windows. Their realised benchmark-relative drawdown distributions should then be judged against the (1 − y)^(1/R) law. Until that test exists, the paper offers clean geometry for stating a leverage budget, while the full allocation depends on drift estimated to within about five points a year.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading day t with a complete 252-session paired history:
    1. Compute SPY and GLD adjusted-close log returns through t.
    2. Estimate annualized log drift:
           mu = 252 * rolling_mean(log_returns)
    3. Estimate and regularize annualized covariance:
           V_sample = 252 * sample_covariance(log_returns)
           V_hat = 0.8 * V_sample + 0.2 * diag(diag(V_sample)) + 0.000001 * I
    4. Convert to estimated arithmetic excess rates:
           b = mu + 0.5 * diag(V_hat)
    5. Solve the intended constrained Kelly problem for w_star using b and V_hat,
       subject to the selected gross-leverage scenario, annualized volatility cap,
       and numerical per-asset bounds.
    6. For each risk target R in {1/7, 1/3, 3/5}:
           k = 2R / (1 + R) in {1/4, 1/2, 3/4}
           u = [SPY: 1, GLD: 0]
           active = k * (w_star - u)
           target = u + active
           cash = 1 - sum(target risky weights)
    7. Reject the entire rebalance if any active weight exceeds 0.50 in absolute
       value, gross leverage or ex-ante volatility is infeasible, required turnover
       exceeds 1.0, required data are missing, or execution costs would exhaust wealth.
    8. Otherwise rebalance at the platform's market-on-close auction convention.
       Subsequent daily targets replace prior targets; a skipped rebalance retains the
       pretrade portfolio. Bankruptcy at an observed execution or valuation time is
       absorbing and terminates trading.