Peng and Wang have a strong monthly US stock sort. A desk could take the one-layer version and leave much of the deep-network work behind.
The trade underneath
At each month-end, the model uses 153 firm characteristics from the Jensen-Kelly-Pedersen panel and 8 Welch-Goyal macro series to forecast next-month excess returns for roughly 5,000 US common stocks. The characteristics are median-imputed, then rank-normalized within each cross-section. Stocks enter deciles by forecast return. The book buys the top decile and shorts the bottom one, value-weights both legs, and rebalances monthly. Training begins with 1963 to 1986 data. The model refits every January on an expanding window, and each forecast averages ten random seeds. The out-of-sample record covers 443 months, from February 1987 to December 2023.
The argument for depth rests on what a residual block can do: it learns a correction while passing its input through unchanged. Under an idealized residual recursion, setting every correction to zero lets a deeper network reproduce a shallower one. "Shallow models are special cases of deep models," the authors write. A plain feedforward stack lacks that pass-through. Gu et al. (2020) found that feedforward nets peaked at shallow depth; Peng and Wang attribute that result to the missing pass-through.
Their comparisons use two pairs. NN is the Gu et al. feedforward net. ResNet adds identity shortcuts to NN. NN+ is the authors' tailored feedforward version, with LayerNorm, LeakyReLU, dropout and AdamW with an l2 penalty. ResNet+ adds pre-activation residual blocks to NN+, with projection paths wherever layer width changes. Four width schedules run from [32] to [32,16,8,4], across depths 1 through 20. Each family has 74 specs. The authors average Sharpe ratios equally within shallow (depth 1 to 5), medium (6 to 10) and deep (11 to 20) groups. On the value-weighted long-short book, deep ResNet+ averages a gross Sharpe of 2.07, versus 1.92 for shallow ResNet+ and 0.89 for deep NN+. Deep ResNet+ earns 3.48% a month with a 5.82% monthly SD; its Fama-French five-factor alpha is 3.35% a month.
Depth one already works
Tailoring delivers the large improvement over the Gu et al. baseline before residual blocks enter the model. Shallow NN earns 0.94 Sharpe; shallow NN+ earns 1.84. Moving ResNet+ from shallow to deep adds 0.15 (deep minus shallow, significant at 5%).
At minimum depth, the residual and feedforward models coincide. Their shared [32], depth 1 cell earns 2.11 Sharpe and has an alpha t-stat of 7.77. It beats the deep-group average. Still, that cell was selected after seeing the grid and is being compared with a 40-spec average, exactly the sort of ex post selection the paper's design seeks to avoid. Narrow bottlenecks also weigh on the shallow average. ResNet+ at [32,16,8,4], depth 5, earns 1.14 against the deep group's 2.07. Much of the measured gain from depth comes from residual blocks repairing a cramped architecture. The one-layer net had no such problem to repair.
What drives the 1.18 Sharpe lead over NN+?
Deep ResNet+ beats deep NN+ by 1.18 Sharpe (significant at 1%) and wins all 40 deep specs. NN+ accounts for nearly the entire gap by deteriorating with depth. Its deep-minus-shallow change is -0.94. Its raw high-minus-low forecast spread contracts from 3.37 to 1.42, while ResNet+ expands from 4.08 to 4.63. Across many deep [32,16,8,4] specs, NN+ Sharpe approaches zero. Plain NN at [32], depth 7 produces a constant forecast in all 443 months, with forecast SD of zero; ResNet at that spec earns 0.90. The authors say their diagnostics "do not identify the cause of forecast collapse."
An optimization failure is the plausible reading. Early stopping checks training loss on the most recent five years of the training sample. We did not find a held-out validation split that might catch a fit as its signal drains away. In the full value-weighted universe, the ResNet minus NN gap is insignificant in every depth group (deep: 0.93 vs 0.82), even after the collapsed NN specs are dropped from that average.
Residual learning acts as insurance against this training pathology. The authors describe it as something that "provides protection against deterioration with depth" and helps in "preventing the collapse of the deep model." They also argue in those passages that depth refines ResNet+. Their evidence is the 0.15 deep-minus-shallow gain (5%), alongside revision accuracy of 1.12% and revision value of 4.92 bps a month, both significant at 1%. Those results hold up. They are small beside the 0.90 delivered by tailoring at depth one, though, and a desk using one hidden layer has no need to pay for the insurance.
The extreme deciles
Deep ResNet+ continues to widen its forecast spread in the tails as NN+'s spread shrinks.
The authors test how the deep forecast reorders every pair of stocks relative to the minimum-depth anchor. Preservation measures the share of pair orderings left alone; revision accuracy checks the changed orderings against realized returns. Deep ResNet+ preserves 0.79 of pairs. Its revisions are right on net by 1.12% and add 4.92 bps a month of revision value. Deep NN+ preserves 0.64; its revisions are wrong on net by 1.88% and cost 34.45 bps a month. The residual net reshuffles about one pair in five and gets the changed calls slightly more right. The feedforward net reshuffles more than a third, gets slightly more wrong than right (-1.88%), and loses 34.45 bps a month on those calls. Zeroing the Size theme takes away 6.11 bps of ResNet+ revision value, the largest single theme effect in deep ResNet+.
Can 260% monthly turnover leave enough?
Deep ResNet+ turns over 259.49% a month. Charging 10 bps per dollar traded cuts its Sharpe to 1.92. At 50 bps, Sharpe falls to 1.30, while deep NN+ reaches -0.26. Those are the only frictions charged: two linear rates, with neither borrow fees nor price impact modeled.
The size screens locate much of the edge. Removing the largest 20% of stocks raises deep ResNet+ Sharpe to 2.78; removing the smallest 20% lowers it to 1.87. The edge is bigger in smaller stocks, where 50 bps looks optimistic. Under the two screens, the deep ResNet+ minus NN+ gaps remain 1.08 and 1.67, both significant at 1%. Averaged over all its width and depth specs, ResNet+ lost 5.15% a month over March and April 2020; NN+ lost 1.57%. We found no drawdown statistics anywhere in the paper.
Peng and Wang show that residual blocks allow a tailored return network to grow deep while retaining its ranking. The after-cost case for depth appears only at the group level. Deep ResNet+ nets 1.92 at 10 bps and 1.30 at 50 bps, against 1.77 and 1.18 for shallow ResNet+, so the net comparison favors depth. The paper gives no per-spec net Sharpe. The [32], depth-1 cell therefore never appears net of costs.
Depth buys about 0.12 to 0.15 of net Sharpe here, a small gain beside the 0.90 that tailoring bought gross.