AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

City angles filter redundant signals; length certificates fail at 0.5

A 60-degree cutoff clears 0.5 history correlation for 94.3% of accepted pairs, against 85.0% unscreened.

2026-10-06 · 6 min read · Signal selection and portfolio construction · US equities

Reviewing: Cities of Signals: Compression, Separation, and the Geometry of Novelty · Marc da Costa Nunes · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Point-in-Time City-Geometry Stock Forecast Selection

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (2)

  • Paper's empirical universe: 20 demeaned assets and 51,833 common stored rows in 76 month blocks; its export lacks verified per-row timestamps.: Use a screened US-stock cross-section, timestamped daily bars and a rolling completed-target window. (invalidates: Direct transfer of the paper's empirical pair frequencies, retained-length distribution and block-certification rates to this stock library.)
  • Paper specifies retrospective screening geometry, not stock forecasts, a monthly training window, portfolio weights, execution timing, turnover, or trading costs.: Prespecify the candidate library, point-in-time selection window, daily long-only ranking portfolio and explicit proportional cost model. (invalidates: Any claim that the paper predicts this portfolio's out-of-sample returns, turnover or net profitability.)

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 75.9%Jul 2024
Sharpe
0.68
Total Return
75.9%
Max Drawdown
-48.2%
CAGR
13.4%
Volatility
24.9%
Beta vs SPY
0.85
Trades
7,032

Nunes's city angle is useful for screening redundant signals in his library. It cannot certify novelty at the commonly used 0.5 threshold, even when paired with retained lengths. The geometry explains why; the observed library explains why the screen works anyway. I would use the angle to narrow a search, then check the surviving pairs directly.

The information a city loses

At each date, a signal forecasts returns across d assets. After demeaning and normalization, that forecast is a point on a sphere. Nunes rotates each date's frame until the realized target points at a fixed North Pole. Every signal receives the same rotation, preserving pairwise correlations; its coordinate along the pole becomes that date's IC. The direction of a signal's time-averaged vector is its city. The vector's magnitude is its retained length.

Full-history correlation D has an exact split: r_i r_j C is the contribution visible through the cities, while K captures how the signals move around their respective means. The cities discard K. Nunes bounds its size by sqrt((1 - r_i^2)(1 - r_j^2)) and constructs cases that reach the bound. The warning is concrete. Two signals can have cities 120° apart, retained length 0.1 and IC 0.05 on every date, yet their history correlation is 0.985. Their shared oscillation vanishes from both averages; normalization magnifies the small remaining difference into a large city angle.

His data come from AlphaNova's Competition May 2026. The main sample contains 776 distinct signals on 20 assets, with 51,833 common rows arranged in 76 month blocks and 300,700 signal pairs. The paper calls the period a multi-year historical evaluation and supplies no per-row timestamps. For the 266,815 pairs whose members both have positive average IC, the history angle equals or exceeds the city angle 58.9% of the time. Within that group, the rate is 74.1% for city angles at most 90° and 1.5% for wider angles.

Why does the uniform null miss this library?

Nunes uses independent, uniformly random histories as a reference. In that model, acute pairs preserve their separation with probability tending to one. Selecting histories with positive average IC gives an overall limit of about 56.0% for 20 assets. He presents the null as a reference for a real library before showing the data, although the abstract still foregrounds limits the library does not follow. The observed 58.9% might appear to support it. Nunes traces the resemblance to offsetting errors: 79.1% of positive-IC pairs have acute cities, versus 56.0% under the null, while their preservation rate is 74.1%, versus a limit of one.

Serial dependence cannot account for that gap within the null's family of independent, rotationally symmetric histories. Estimated lag profiles put the median pair dependence factor at 2.4 and the acute-pair lower bound at 98.06%. Reaching 74.1% would require a factor near 390. Another diagnostic gives 19 times the squared mean of D as 1.076, above the cap of one on its expected value under the null with any temporal dependence. Nunes treats this as descriptive, not as a rejection test. Positive residual co-movement appears in 91.0% of positive-IC pairs.

The paper's own warning is direct: "These frequencies describe the observed library; they do not validate the uniform null." Nunes instead derives a stationary ergodic margin, g = (1 - ρ_X ρ_Y)C* - κ, that allows signals to move together. Its theorem requires stable, nonzero population means. Here the median retained length is 0.0338, overall preservation ranges from about 42% to 64% across time windows, and the last month alone records 38.9%. Nunes acknowledges that stability has not been established. The mathematical result stands; the empirical evidence remains descriptive.

A 60-degree cutoff earns a look

Across all 776 signals, a city correlation cutoff of at most 0.5 accepts 217,926 of 300,700 pairs, or 72.5%. History correlation is at most 0.5 for 94.3% of those accepted pairs. The unscreened rate is 85.0%, so the share above 0.5 drops from 15.0% to 5.7%.

The screen misses some pairs in both directions. It accepts 27.5% of pairs with D above 0.5 and discards 19.6% of acceptable pairs. Setting the cutoff to zero raises success to 96.8%, but retains only 23.3% of all pairs; acceptable-pair retention falls from 80.4% to 26.5%. The 9.3-point gain survives resampling, with ranges of 8.2 to 10.0 points by month blocks, 7.2 to 12.1 by scientist clusters, and 7.1 to 12.2 using both.

The mean term can contribute at most 0.0659 to any pair's D. Even so, city correlation ranks history correlation with Spearman 0.561. In this library, city alignment tracks residual co-movement, which the geometry itself cannot promise elsewhere. These comparisons also use signed correlation: a nearly opposite signal counts as novel. The paper flags the absolute screening rule |D| ≤ τ as unevaluated.

Certification needs much finer blocks

Adding retained lengths makes a deterministic certificate possible. At 0.5, it certifies zero pairs in this sample. Even the largest retained length, 0.2567, produces an upper endpoint of at least 0.8682. Nunes then stores one mean vector per time block to recover information lost in the overall mean. With 64 blocks or fewer, certification remains at zero. Month blocks certify under 0.05% of acceptable pairs. It reaches 54.7% at 16,384 blocks, roughly a third as many vectors as rows, and 83.4% at 32,768 blocks. The abstract acknowledges the cost: block means recover the missing information "here only at fine resolution". Nunes's proposed storage choice is the mean vector rather than the city alone, followed by finer block means for pairs still undecided. That keeps the inputs to the guarantees explicit.

Coarse blocks weaken screening too. Month-block correlations accept 83.5% of pairs, compared with 72.5% for the single city. They admit 39.8% of pairs with D above 0.5, versus 27.5%, and success falls from 94.3% to 92.8%. Their advantage is retention: they reject 8.8% of acceptable pairs, rather than 19.6%. On this library, substantial compression and useful certification rarely arrive together.

Our US equity run

We built a point-in-time version on US stocks. The figures here are ours, from a backtest covering 2020-01-01 to 2024-07-01. We refreshed a universe of 40 liquid stocks yearly. Eight forecasts competed: two momentum, two reversal, two low-volatility and two fundamental ratios. Each month, we formed the frame from up to 252 dates with already closed targets and retained positive-IC candidates. We accepted up to four, provided each new candidate's city dot product was at most 0.5 against every previously accepted candidate. The combined forecast selected ten equal-weight longs each day.

After commissions of $0.004 a share, our run returned 75.86%. Its Sharpe was 0.68, maximum drawdown 48.16%, and trade count 7,032. Beta to SPY was 0.85, with annualized volatility of 24.93%, giving this long-only book market-like risk. A top-ten long-only portfolio drawn from 40 large names has substantial market exposure, and the window includes 2020 and 2022.

There is no portfolio return in Nunes's paper to compare with ours. His 94.3% measures pair geometry in a 20-asset competition library; our Sharpe measures something else. More damaging for attribution, we built unscreened and direct-correlation arms but do not have their returns separated. Our figures therefore assign no dollar gain to the city screen. Eight candidates create only 28 pairs beside the paper's 300,700, and we did not model market impact. Our run is one automated pass through the trading implementation, not a verdict on Nunes's work or his pair statistics.

The prospective comparison

Nunes acknowledges that city construction uses realized targets; prospective screening performance remains unestablished. The comparison still needed would build city, direct-correlation and unscreened arms from closed targets, give them identical candidate priority, and score next month's realized pairwise D. A long-short, beta-neutral portfolio difference would keep our 0.85 beta from swamping the return comparison. In an earlier note, we argued that signal correlation weakly proxies for PnL dependence. City correlation compresses the signal another step. Until those arms are separated, the 9.3-point lift belongs to AlphaNova's archive.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading year, refresh the 40-stock liquidity screen.
Before the first trading session's close each month:
  Form a common, ordered cross-section on completed diagnostic dates only.
  Use up to 252 completed-target dates; require at least 126.
  Normalize each candidate forecast and subsequent return target by date.
  Retain candidates with positive mean matched-date IC; rank by that IC.
  Accept up to four candidates in priority order if each city dot product
    against an accepted candidate is <= 0.5.
  Record retained-length certificates separately; an inconclusive certificate
    does not alter selection.
Each trading day, combine selected candidates' current normalized raw forecasts,
rank eligible stocks, and target 10% in each of the top ten; leave unused weight
in cash. Execute eligible changes at observed closes.

Matched point-in-time direct-correlation and unscreened selection arms use the same candidate priority and portfolio rules; their comparative returns are not supplied here.