The tradeable result is thin. Lin, Chen, Wang, Briola and Aste set a network's depth, widths and connections before training, using the correlation matrix of firm characteristics to determine its form. The resulting net has 21.0 thousand parameters. Its pooled out-of-sample R2 matches a hand-tuned three-layer benchmark, 0.509% against 0.475%, p = 0.796. NN3 has the lower MAE, 0.1057 against 0.1059, and the larger asset-weighted gross spread, 2.44% against 1.99%.

The stock-selection advantage over that benchmark reaches one metric.

The authors acknowledge as much in the introduction. Following Holm adjustment, only the Pearson IC differences remain. They describe HNN (marginal) and NN3 as "statistically indistinguishable" on R2, Spearman IC and portfolio spread. The abstract still says the models "rank the cross-section more accurately". Their response carries some force. The ranking lead over both structural controls survives a common turnover penalty, with unadjusted p = 0.007 against the shuffled network and p = 0.028 against the dense one. The paper also gives ranking priority over level accuracy because a long-short portfolio earns returns by ordering stocks within each month.

We did not test their method. Our data cannot reconstruct the exact 94-characteristic GKX panel or the 1957 to 2016 sample. A substitute panel available to us would run mostly from 2010 onward, while the clique filter, network mapping and walk-forward training would require fresh implementation. Nothing below tests their claim.

The architecture

The paper's wager is that return-relevant interactions among firm characteristics are sparse and genuine. Its sparse network retains 21.0k of the 1,680.7k weights in a width-matched dense network, about 1.3%. The authors estimate the dependence structure up front and use it to choose the form, instead of searching over depth and width.

Each month, they rank characteristics cross-sectionally into [-1,1] and replace missing values with zero. They then calculate the absolute Pearson correlation matrix across characteristics. A Maximally Filtered Clique Forest processes that matrix, retaining the strongest dependence while imposing a decomposable graph with cliques capped at size K. Those retained subsets become neural units. Layer 2 contains the surviving pairs, layer 3 the surviving triples, and higher layers follow the same scheme. Set inclusion determines the connections, giving every order k unit exactly k incoming weights. A linear readout combines every non-input layer into the forecast. This lets a unit that belongs to no larger clique reach the output. The only structural choice is K, searched over {2,...,7}. In the paper's words, model selection "is not eliminated but changes in kind."

The GKX panel contains 94 firm characteristics and monthly returns from 1957 to 2016. The target is next-month excess return over the one-month risk-free rate. Training begins in 1957 and expands annually. The following twelve years form the validation period, followed by the next year as the test period. Testing starts in 1987 and ends in 2016, yielding 30 non-overlapping test years and 2,508,749 test observations. For each window, the correlation matrix uses 150,000 randomly sampled training observations. A single graph then applies across every firm and date in that window.

The inputs exclude industry dummies and macro interactions, as the authors flag. They frame the exercise as a controlled architecture comparison instead of a replication of GKX's 920-input specification. The benchmarks are Huber-3 (size, book-to-market, 12-month momentum under Huber loss), PCR on all 94, and NN3 at 32-16-8. Results for each neural model average a ten-seed ensemble.

The headline figures favour HNN on several columns. Pooled out-of-sample R2 against a zero forecast is 0.509% for the characteristics-only model, HNN (marginal), versus NN3's 0.475%. HNN (m-s) reaches 0.511%. That version takes the union of cliques estimated above and below the training-median excess return. Pearson IC is 0.0642 versus 0.0586. The equal-weighted decile long-short spread is 3.95% per month gross and 3.34% net at 25 basis points per traded dollar, with an annualised Sharpe of 2.28 across 359 rebalances.

Does the ranking advantage reach P&L?

Only one of six measures survives. The authors compare squared error, absolute error, both ICs and both decile spreads with three alternatives, then apply Holm adjustment to the resulting 18 p-values. The Pearson IC improvement of 0.0056 has p = 0.002 unadjusted and remains significant after adjustment at p = 0.034. The R2 difference is distant at p = 0.796. Spearman IC is 0.0517 against 0.0474, with p = 0.178. NN3 records a slightly lower MAE, 0.1057 against 0.1059, p = 0.050. At 25bps, the net equal-weighted spread comparison is 3.34% against 3.13%, p = 0.112.

The authors' defence returns to portfolio mechanics: long-short returns depend on ordering stocks correctly within a month. Yet their own Huber-3 result complicates that argument. Huber-3 has the highest Spearman IC in the table at 0.0676, alongside the lowest Pearson IC at 0.0227 and the lowest equal-weighted spread at 0.51% per month. The paper uses this row to support reporting both coefficients. It also shows how little statistical-to-P&L support a Pearson-only victory provides when the portfolio test does not clear.

The shuffle carries the evidence

MLP-HNN preserves the induced widths and makes the connections dense. Parameter count rises to 1,680.7k from 21.0k. Its Pearson IC is 0.0524, its Spearman IC is 0.0436, and its equal-weighted spread is 3.47%, all below the sparse architecture. The rest of the row matters. Pooled R2 is 0.429%, the weakest of the four neural specifications, while MAE is slightly lower at 0.1057. Its asset-weighted gross spread reaches 2.48%, the table's highest and above NN3's 2.44%. Net performance is 1.85% at 25bps, with a Sharpe of 0.99. As the authors observe, this comparison changes both the connectivity and the readout head. It therefore measures complete architectures rather than sparsity alone.

The shuffled model gives the cleaner comparison. Capacity, topology and parameter count stay fixed, leaving only the node assignment mismatch. It uses the same graph, widths and 21.0k parameters, with characteristics permuted across nodes. Pearson IC declines from 0.0642 to 0.0597 and survives Holm at p = 0.003. R2 moves from 0.509% to 0.496%. The equal-weighted spread drops from 3.95% to 3.73%, p = 0.007 unadjusted, then disappears after adjustment. Compare that outcome with the abstract's statement that, for the two ablations, "both effects remain significant after correcting for multiple testing". The statement holds only for Pearson IC. The specific clique assignment contains information.

Parameter economy is measured against the dense comparator. NN3 uses 3.8k parameters beside HNN (marginal)'s 21.0k, roughly a fifth. HNN (m-s) selects K <= 4 in 20 of 30 annual refits and has a median parameter count of 5.2k, yet it does not improve the ranking result over the characteristics-only model.

Small stocks drive the spread

The 3.95% result is concentrated there. The equal-weighted gross spread is 3.95%, while the asset-weighted gross spread falls to 1.99%, below NN3's 2.44%. After 25bps, the asset-weighted figures are 1.31% for HNN (Sharpe 0.73) and 1.78% for NN3 (Sharpe 0.95), with p = 0.036 in NN3's favour. Break-even cost is 159bps under equal weighting and 73bps under asset weighting, compared with NN3's 153 and 91. Mean one-way turnover across the two unit-notional legs is 1.25 and 1.36.

A Sharpe of 2.28 in this setup is not a number anyone trades. The paper makes the drivers plain: equal weights, full-universe deciles, no market impact or borrow costs, and a break-even level of 159bps against the 25bps charge used in the test. The authors describe the cost calculation explicitly. A single uniform linear rate is applied to measured turnover. Security-specific spreads, short-borrow fees, market impact and capacity are omitted. The gap between their own equal-weighted and asset-weighted results, 3.95% gross against 1.99%, measures the exposure to that omission.

After 2007

Both neural models trail a zero forecast during 2007 to 2016. Pooled R2 is -0.135% for HNN (marginal) and -0.138% for NN3. HNN (m-s) leads the neural models in that period at -0.040%, while PCR alone is positive at 0.087%. HNN (marginal) still produces a monthly spread of 2.55% over the decade, against 2.48% for NN3. The paper states the split precisely: "Level accuracy and cross-sectional ordering therefore come apart in the later sample." In annual comparisons, HNN (marginal) exceeds NN3 in 19 of 30 years on R2 and 19 of 30 on Pearson IC. The paper does not test that lead for significance. Its sample finishes in 2016.

Selection deserves attention as well. K and the learning rate are selected through validation MSE once every five test years, then retained for the block. Graphs and weights are fitted again each year. The validation window consists of the twelve years directly preceding the test year, which means later windows overlap previous test years. No future data enters. Architecture selection remains coarse and carries sampling noise from the 150,000-observation correlation estimate.

A useful prior, with a narrow win

Replacing a grid over depth and width with one bound on interaction order is a sensible default for characteristic-panel forecasts. HNN (marginal) uses 21.0k parameters, or 1/80th of the width-matched dense network's 1,680.7k. The dense model also trails on both IC measures and the equal-weighted spread.

Production use would require an asset-weighted or liquidity-capped implementation in which the Pearson IC advantage appears in a spread that passes its own test. In the reported results, asset weighting returns the lead to NN3: 1.31% against 1.78% net at 25bps, with p = 0.036 pointing the wrong way.