AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

A continuous return chart without a ranking test

Liu captures entry-date sensitivity, while his fund-picking criterion lacks an effective sample size

2026-09-09 · 7 min read · US equities and US ETFs

Reviewing: From Discrete Trailing Returns to a Continuous Graphical Profile: Return-to-Present Curves · Lei Liu · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Fixed-Endpoint RTP Persistence Momentum versus SPY

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

Run on a different market than the paper

The illustrations include retirement mutual funds, which are not an available holdable universe. Implement the same fixed-endpoint return-profile mechanism on US stocks and ETFs, where it remains valid because RTP only requires adjusted historical price series; any results would apply to the adapted stock/ETF universe rather than the paper's fund examples.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully

  • deviation left undescribed by the audit (invalidates: The paper's non-predictive interpretation does not establish any expected trading performance for this extension; none of the paper's figure-specific descriptive predictions constitute prospective performance targets.)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 53.4%Jul 2024
Sharpe
0.43
Total Return
53.4%
Max Drawdown
-59.1%
CAGR
10.0%
Volatility
27.5%
Beta vs SPY
0.92
Trades
2,101

A fund ranking that changes between standard look-back windows deserves to be seen before anyone trades it. Liu's fixed-endpoint chart does exactly that, and the NVDA versus AMD example makes the case cleanly.

Fund pages usually report 1-month, 3-month, 6-month, 1-year, 3-year and 5-year returns. Liu treats those six observations as samples from a continuous function. Between them, the ordering of two assets can reverse.

The abstract already limits the claim. The Return-to-Present curve is intended for "historical comparison and decision support rather than prediction or statistical inference." The narrower dispute comes from the discussion, which still presents a "direct historical criterion for choosing between investments that may otherwise appear comparable".

The chart itself

Start with adjusted prices. At look-back horizon h, define R(h;t) = P(t)/P(t-h) - 1. Hold the evaluation date t fixed, sweep h through the available history, and draw the result. Since every curve equals zero at h = 0, the lines meet at the chart's right edge before spreading left. Liu labels the horizontal axis by purchase date u = t - h, keeping the display in calendar time and allowing actual transaction dates to be marked.

There is no additional return measure. Liu describes the contribution as "representational rather than algebraic," adding that "the principal contribution of RTP is visual rather than computational." He also plots D_AB(h;t) = R_A - R_B against a selected reference asset. Positive values favour A, negative values favour B, and each zero crossing marks an entry date at which the ranking reversed.

For a trader, the chart separates relative performance from sensitivity to the chosen start date. A difference curve that keeps the same sign across five years of entry dates gives a ranking that survives changes in the chart origin. Crossings show that the start date drove the answer.

Four pictures of past winners

The illustrations compare NVDA with AMD, three AI-focused ETFs with different inception dates, three mutual funds offered through the Washington University in St. Louis retirement plan, and a hypothetical move from DIA into QQQ.

The semiconductor comparison works best. For entry dates roughly three to five years before the endpoint, the NVDA-minus-AMD difference reaches several hundred percentage points. At some dates it exceeds 1,000 percentage points. The curve then crosses approximately three years before the endpoint. AMD leads through a substantial part of the later period, reaching its largest advantage around 500 days before the evaluation date. A table containing 1-month, 3-month, 6-month, 1-year, 3-year and 5-year returns would miss the location of that crossover.

A full sign flip within one sector pair.

The ETF panel is the feature worth keeping. AIQ launched May 11, 2018, CHAT May 18, 2023, and AIS December 3, 2024. Starting a conventional common-origin chart at the newest inception removes more than six years of AIQ history and approximately eighteen months of extra CHAT history. Placing since-inception figures beside one another creates another problem: the numbers cover different calendar periods and holding horizons. Liu instead shortens the left side of each curve and calculates differences only where the histories overlap. Unequal histories appear as unequal line lengths, without requiring an analyst to select a shared starting date.

The rotation example has an investor sell DIA and reinvest in QQQ on January 2, 2026. Both choices are then evaluated at the common endpoint August 20, 2026. The DIA-minus-QQQ difference at entry is about -5 percentage points, leaving QQQ's RTP about five points higher. A staggered transaction, selling January 2 and buying January 3, can be read from the same curves because they share the endpoint. Verification requires a price series extending through August 2026. A genuine DIA-to-QQQ trade in a taxable account would also incur a spread and a realised gain, neither included in the 5 points.

Liu frames the retirement-fund panel around recurring contributions, which arrive at multiple purchase dates. Each point still represents the return from one purchase on a candidate contribution date. An actual contribution stream has its own internal rate of return, and the figure does not supply it. The panel shows VPMAX ahead of JLGMX and VIIIX for most historical entry dates in a five-year window. Their relative order looks stable, although the magnitudes are not reported.

The endpoint remains a choice

RTP shifts the arbitrary date from the beginning of the chart to the end. Liu acknowledges the consequence. He writes that a large price move near the endpoint affects many look-back horizons, allowing an exceptional recent episode to influence much of the curve. Alternative endpoints "should therefore be substantively motivated rather than selected to favor a particular result," while sensitivity to exceptional endpoint-adjacent periods "warrants further study."

His stated preference is the present as the natural endpoint for a current comparison. A historical endpoint can make sense when tied to a particular market, company, or policy event. In live use, the present supplies the date. In a published figure, the author selects it and the reader must accept the stated motivation.

The mechanical effect falls mainly on the level curves. One large endpoint move in a single name shifts its entire curve and, by Liu's account, changes returns across many look-back horizons simultaneously. Common sector moves partly cancel in D_AB. His case for comparisons among same-sector peers is persuasive on that point.

A criterion with no effective sample size

The discussion suggests combining the proportion of entry dates favouring each investment with the magnitude and persistence of RTP differences. Liu calls the result a descriptive summary. The qualification arrives immediately: "This does not establish future superiority, but it provides a direct historical criterion for choosing between investments that may otherwise appear comparable."

The warning becomes stronger one paragraph later. Adjacent RTP observations "are strongly dependent because they share the same endpoint and largely overlapping price histories." Liu identifies two possible summaries, the proportion of historical entry dates favouring one investment and the number of curve crossings. Both "should not be interpreted as independent evidence or as probabilities of future outperformance." The resulting criterion describes one price path viewed from one endpoint.

Those statements fit together, but they leave the proposed decision rule without an effective sample size. Sweeping a five-year look-back horizon turns the favouring proportion into a shape statistic for a single path. The narrow cross-section adds another weakness. Every illustration uses two or three selected tickers, compared with the 29-stock panel discussed in an earlier note. The NVDA/AMD, ETF and fund figures also lack a stated evaluation date, data source and price frequency, making the drawn examples difficult to verify.

Our stock and ETF implementation

Liu includes retirement mutual funds that we cannot hold. We therefore used the same fixed-endpoint return-profile mechanism for US stocks and ETFs. The substitution is valid because the construction requires only adjusted price series, though every result below belongs solely to that adapted universe. His fund illustrations remain outside it.

We converted the display into a selection rule. Liu proposes no strategy, so the exercise evaluates our rule rather than his work.

At each rebalance, we select the top 500 non-ADR stocks and ETFs by prior-year capitalisation from point-in-time screening data and exclude SPY. Using prices through the previous close, we calculate D_i(h) = R_i(h) - R_SPY(h). For each of four portions, 21, 63, 126 and 252 sessions, we measure the positive-difference share, mean, median and sign-change count. Each measure is percentile-ranked across the universe, then averaged. The portfolio holds the Top 20 names at equal weight with a 10% cap for 20 trading days. It is long only and executes at the closing auction. Every fill incurs commissions of four tenths of a cent per share, subject to a $1 order minimum, before performance is calculated. Window: 2020-01-01 to 2024-07-01.

One automated pass over that window produced total return 53.41%, Sharpe 0.43, Sortino 0.60, volatility 27.50%, max drawdown -59.10% and Calmar 0.17. The result is weak. A 0.43 Sharpe paired with a -59.10% drawdown does not describe a book anyone should run. Liu reports no strategy performance because he offers no strategy, leaving no like-for-like figure to compare with ours.

These figures speak first to our implementation. Within each portion, the four features come from the same overlapping difference curve and therefore act as dependent readings of one name. Liu's caution about dependence applies directly to our score. In practice, the composite resembles a blend of trailing relative returns over 21, 63, 126 and 252 sessions.

The portfolio also holds 20 large caps, long only and unhedged, through the March 2020 crash and the 2022 drawdown. Realised volatility reached 27.50%. Max drawdown was -59.10%, with the market path accounting for much of the equity curve.

What would change the verdict?

RTP deserves a place as a chart. Its treatment of unequal inception dates is the part I would want on a platform tomorrow. It retains more than six years of AIQ data and about eighteen months of CHAT history that disappear from a common-origin overlay.

Liu writes that whether RTP's historical patterns predict future returns is "a separate empirical question," and says RTP "is therefore not proposed as a new momentum factor or trading signal." Yet the discussion gives readers a fund-selection criterion based on summaries that the same discussion says should not be treated as independent evidence or probabilities of future outperformance. The criterion remains unsupported.

A universe-wide test showing that the persistence features add to a plain 12-month-minus-1-month rank, using more than two or three names in each comparison, would change my view of the paper's second half.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each scheduled 20-trading-day rebalance date t:
  1. Resolve the point-in-time universe from prior-calendar-year screening rows.
     Keep non-ADR STOCK and ETF records and take the top 500 by capitalization.
     Exclude SPY from candidate selection.
  2. Set the signal cutoff to the preceding trading-day close.
  3. Require at least 253 matched observations and complete candidate/SPY
     histories for the 21-, 63-, 126-, and 252-session portions.
  4. For each candidate i and each historical offset h, calculate:
       R_i(h)   = P_i(t) / P_i(t-h) - 1
       R_SPY(h) = P_SPY(t) / P_SPY(t-h) - 1
       D_i(h)   = R_i(h) - R_SPY(h)
     Retain h=0 in the curve but exclude it from summaries because D_i(0)=0.
  5. Within each portion, calculate positive-difference share, mean difference,
     median difference, and adjacent-sign zero-crossing count.
  6. Cross-sectionally percentile-rank the first three features descending and
     crossing count ascending. Equal-weight the four ranks to form each portion score.
  7. Equal-weight the four portion scores and rank candidates descending.
  8. At the scheduled close, hold the top 20 eligible candidates at equal weights,
     subject to a 10% position cap. Use observed close prices only; skip trades whose
     required execution close is missing.
  9. At the next scheduled rebalance, exit names outside the top 20 or absent from
     the current screened universe and rebalance retained and new positions.

Performance metrics are calculated after the platform's per-fill cost recalculation.