Change the learner and the change-in-inventory coefficient jumps from -0.037 to -0.387. That spread comes from the headline empirical panel. I would take the econometrics home and leave the asset-pricing conclusions at the door.

Remove the factors before fitting the learner

Eighty-nine lagged characteristics are set against 72 months of returns. Because returns load on a few latent common factors with heterogeneous exposures, unit and time dummies cannot clean the data. The required term is interactive fixed effects, λ_i'f_t. Bai's (2009) standard IFE estimator assumes that observed controls enter linearly and remain few relative to N and T. This application satisfies neither condition.

Chen, Polselli and Clarke handle the two problems separately. They begin by projecting out the factors using cross-sectional averages of the covariates, following the spirit of Pesaran's CCE. For the p>T case, they extend the projection by taking eigenvectors from the covariance of those averages, following Rücker, Vogt, Linton and Walsh. An eigenvalue threshold determines the factor count: retain eigenvalues at least α times the largest, where the analyst chooses a small value such as 0.01 or 0.05.

Chernozhukov et al.'s double machine learning then runs on the projected data. Lasso, gradient boosting or a feed-forward net learns the nuisance functions l0 and m0 from held-out folds. The orthogonal score multiplies the two projected residuals. Standard errors are clustered by firm.

If the procedure works as advertised, it offers a clean estimate of whether a characteristic moves next month's excess return after removing latent factor exposure and nonlinear confounding from the 89 other lagged characteristics. For a researcher, that makes it a signal-validation instrument.

The evidence has two parts. The Monte Carlo uses 100 replications in each cell, θ0 = 1, two factors and only two relevant covariates among p. N ranges from 20 to 500, T lies in {30, 50, 100}, and p lies in {50, 100, 300}, with five-fold cross-fitting. The empirical application covers 29 current or recent large-cap Dow constituents across 72 monthly periods, from April 2016 through March 2022. It tests seven candidate treatments against 89 lagged characteristics from the Green, Hand and Zhang set, using threefold block cross-fitting and 100 random-search draws for each learner.

How much do the simulations establish?

The feasible-versus-infeasible bias decomposition is persuasive. In the discontinuous design at N=20, the linear IFE estimator has feasible bias above 0.20 in nearly every cell, including 0.2446 at T=30/p=50 and 0.2947 at T=30/p=300. Give the infeasible version the true factors and its bias remains around 0.20, locating the problem in misspecification. A larger cross-section barely changes it. At N=500, T=100, p=50, IFE bias is 0.1989, compared with 0.2066 at N=20. Boosting records bias of 0.0384 at N=20, T=100, p=50 and 0.0716 at N=500 with the same T and p.

Practitioners should pay closer attention to the standard errors. Under the linear design at N=20, T=30, p=50, the average analytic IFE standard error reaches 1.6615 while the empirical SD is 0.0711. At p=300, those figures are 3.4808 and 0.0989. With wide control sets, IFE inference becomes conservative by one to two orders of magnitude because the estimator screens out nothing. Anyone using linear IFE t-stats to screen characteristics may therefore find nothing for a mechanical reason.

Bias falls with T. Raising N shrinks variance.

For DML-Lasso under the discontinuous design at p=300, bias declines from 0.0630 at T=30 to 0.0326 at T=100 while N stays at 20. Moving N to 500 mainly reduces variance. The paper describes the same pattern, saying bias "is reduced primarily by larger T rather than larger N". Equity researchers usually feel the stronger pull toward a wider universe.

The gain has a narrow scope. Footnote 7 says the (N, T, p) grid was calibrated to the application at (29, 72, 89), meaning the simulation examines this panel, and only this panel. Reliability is argued from SE/SD ratios near unity; we did not find reported coverage or rejection frequencies. The authors identify one possible failure themselves. In larger, higher-dimensional cells under the discontinuous design, boosting SEs run slightly below the SD, which they say may produce mild over-rejection.

A thin empirical panel

NT is 2,088, while p lies close to T. The CLT used for inference runs across 29 units. Assumption 2.10(c) requires cross-sectional independence among those units. The authors explicitly leave relaxation through a mixing-based CLT to future work. After removing two factors, we would still expect dependence among 29 mega-caps. Data-driven factor selection also appears on their future-work list, and we did not find the α used for the application.

The universe fixes membership at the sample's end, defining it from current or recent Dow constituents.

Learner disagreement then takes over. Change in inventory is insignificant under OLS at -0.144, SE 0.299, and under IFE at -0.07, SE 0.159. DML-FE allows flexible controls while keeping fixed effects additive. Its estimates are -0.107 (0.081), -0.032 (0.084) and 0.041 (0.095), with the last reversing sign. Every DML-IFE learner reports significance, though at different levels: Lasso gives -0.037 (0.013) at 1%, boosting gives -0.387 (0.234) at 10%, and NNet gives -0.195 (0.097) at 5%.

The largest coefficient carries the weakest evidence. Magnitudes differ tenfold, and sign is their sole point of agreement. The paper interprets this row as an effect recovered by the framework and missed elsewhere. It also notes that higher RMSE_m under Lasso suggests trouble partialling out the treatment in that panel. A shared sign still supplies no position size.

Industry-adjusted size produces direct sign disagreement. Lasso reports 0.000 with p<0.01, whereas NNet reports -0.002 (0.001), also at p<0.01. The Lasso fit has RMSE_m of 24.706. The authors themselves call that estimate unreliable and treat the panel as the exception to their account. The table contains seven treatments and eight estimator columns, and we did not find a multiple-testing adjustment.

The paper summarizes the table in two opposing ways. Its abstract says "several effects documented under linear specifications lose statistical significance once high-dimensional nonlinear confounding and the presence of IFE are jointly accounted for". The introduction foregrounds the reverse claim: "panel DML-IFE recovers significant effects of inventory changes and of scaled earnings forecasts on excess returns that IFE fails to detect." Rows support both descriptions. The method's apparent purpose depends on which quotation leads.

Two of the eight cells in the turnover volatility row lack stars, despite the paper's summary that the effect is significant across all estimators and learners. DML-FE NNet prints -0.004 with a standard error of 0.003. DML-IFE NNet prints -0.001 (0.002).

The projection removes what factor investors collect

The bid-ask spread row makes the issue plain. OLS estimates 0.946 (0.222, p<0.01), while IFE estimates -0.207 (0.343). DML-FE, flexible in the controls and additive in the fixed effects, disagrees across learners: 0.161 (0.042, p<0.01) for Lasso, 0.941 (0.46, p<0.05) for boosting and 0.04 (0.03) for NNet. Every DML-IFE estimate is approximately zero: -0.007 (0.084), -0.159 (0.28) and -0.006 (0.344). The factor projection eliminates the liquidity premium. The paper reaches the same interpretation, calling the OLS premium "attributable to unmodelled interactive fixed effects rather than a genuine causal relationship".

That creates the trading problem. Π̂ removes a characteristic whose return comes through exposure to a pervasive common factor before forming the score. The estimand captures the partial effect after common factor exposure has been removed. A long-short book earns that exposure. The paper answers the causal question openly, while readers of this site are likely seeking the tradable one.

Breadth binds first. Twenty-nine names are rebalanced monthly, and the coefficient ranges from -0.037 to -0.387 depending on the learner. Cross-sectional averages over the full 72 months also enter the projection. This is acceptable for estimation, though trading would require a recursive version.

We judged the method buildable and are running a version now. The paper and our implementation both use U.S. equities, so there is no substitution across asset classes. We cannot reproduce the exact Dow constituent list point-in-time and instead use a top-N large-cap screen as a liquid Dow-like proxy. That substitution is ours.

A wider panel would change my view. Apply the same seven treatments to a few hundred names over twenty years, refitting the projection recursively. If Lasso and boosting then keep the coefficients within a factor of two, I will believe the empirical table. Since the simulation shows that bias falls with T, a long sample is the cheap experiment. Until someone runs it, the paper offers a well-built estimator with a demonstration attached. I would push back on the conclusion's claim that the application "further demonstrated the method's practical relevance".