Small-cap ESG remained exposed to financing conditions throughout 2021-2024. Kocaarslan's clearest differentiated finding rests on two variables, the S&P 500 Bond Index and the BAA-minus-AAA default spread. When the model explains the S&P SmallCap 600 ESG index, those variables interact most often with the S&P SmallCap 600 benchmark. A trading desk should read that as evidence that small-cap ESG exposure moved alongside credit conditions during the sample.

The paper studies three targets: the S&P 500 ESG, S&P MidCap 400 ESG and S&P SmallCap 600 ESG indices. Its daily sample begins on January 11, 2021 and ends on November 22, 2024. Each target level is predicted from the fifteen market, risk and uncertainty series listed in Table 2. The inputs cover the three conventional size indices, bond and commodity indices, policy rates, spreads and five implied-volatility measures.

Four tree models enter the comparison: CatBoost, LightGBM, XGBoost and random forest. The sample is divided into 75% training and 25% testing sets, with hyperparameters selected through cross-validation. TreeSHAP allocates each prediction across the individual inputs after the winning model has been chosen. XGBoost wins for the large-cap and mid-cap targets; CatBoost wins for small-cap.

Across all three ESG targets, the corresponding size benchmark has the strongest positive association.

For large-cap and mid-cap, interactions involving the benchmark remain within the three equity size indices. Small-cap departs from that pattern. Its base index interacts most frequently with the bond index and the default spread.

How much comes from index construction?

Each ESG index is screened and reweighted from its matching S&P size universe. The S&P 500 ESG begins with the S&P 500, the MidCap 400 ESG with the MidCap 400, and the SmallCap 600 ESG with the SmallCap 600. Kocaarslan acknowledges the consequence in a footnote: "a degree of strong comovement between ESG and non-ESG benchmarks is expected ex ante."

The Discussion resists that interpretation of the SHAP rankings. It says that "these results are not merely mechanical or driven by the partial constituent overlap between ESG and benchmark indices." The explanation follows: "ESG screening removes noneligible firms, alters sector exposures, and shifts factor sensitivities, meaning that ESG portfolios often differ meaningfully from the underlying size universe." The constituent sets certainly differ. This specification still leaves the argument exposed.

The models use index levels. Across the best models, test R-squared ranges from 0.99861 to 0.99956, and no non-ESG target is modelled with the same inputs. At those fit levels, the same-day parent index and its ESG child move close to an identity. Variation introduced by screening vanishes into the fourth decimal.

The individual results are equally stark. XGBoost records 0.99956 for LargeCapESG and 0.99872 for MidCapESG. CatBoost reaches 0.99861 for SmallCapESG. Their RMSE values, in the same order, are 1.1471, 1.0492 and 1.1087 index points.

MAPE determines which model wins for each target. On LargeCapESG, XGBoost produces 0.0022670, compared with CatBoost at 0.0028151, LightGBM at 0.0028056 and RF at 0.0026712. The Diebold-Mariano test rejects equal accuracy with p = 0.0056, 0.0034 and 0.0044. SmallCapESG reverses the ranking. CatBoost achieves 0.0021376, ahead of XGBoost at 0.0027232, LightGBM at 0.0031214 and RF at 0.0028085. It defeats LightGBM (p = 0.0000), XGBoost (p = 0.0029) and random forest (p = 0.0000). MidCapESG offers no clear choice between the leading pair, with XGBoost against CatBoost at p = 0.8607.

A 0.99956 R-squared has no trading use

The reported scores measure contemporaneous fit in index levels. Every feature is observed on the same day as the ESG target, and the paper sets the forecasting horizon at h=0. The target day's information must arrive before a position can be formed under this design. There is no return signal, holding period, turnover estimate or cost assumption.

The random 75/25 split distributes observations from 2021-2024 across training and testing. Kocaarslan concedes the resulting regime leakage and describes the metrics as "best viewed as measures of in-sample generalization across heterogeneous states of the world." Diebold-Mariano testing uses h=1 because the test requires h of at least 1, while the operating model remains same-day. Those p-values rank algorithms within the randomized evaluation. They provide no evidence of forward forecasting ability.

Working in levels creates a further judgment call. The authors acknowledge that machine learning models "can still exploit persistent components in the data." Their response is categorical: "the elevated R2 values in our analysis reflect model completeness rather than overfitting or multicollinearity issues." Completeness is the wrong defence here. The target's parent index appears among the predictors, measured that same day and in levels, so the exercise approaches an identity.

A trader seeking an implementable test would need to select returns, changes, residuals from the parent index, or a chronological level model. Every choice poses a different question from the paper's h=0 attribution exercise.

The credit-spread clue

The small-cap interaction still has value for risk monitoring. Over the sample, the BAA-minus-AAA spread averages 0.8698 and runs from 0.61 to 1.23. The federal funds rate rises from 0.05 to 5.33, while VIX spans 11.86 to 38.57. This 3.9-year window contains a post-pandemic recovery and a rapid tightening cycle. It does not span a credit cycle.

SHAP assigns each prediction across per-input contributions. It offers no estimate of a firm's default probability. The paper states the limit directly: the interaction effects "characterize the structure of the predictive function learned by the machine-learning models." Figures 2, 4 and 6 display the interaction rankings. We found no numerical interaction magnitudes and no evidence on stability across the four algorithms.

Interpretation turns on one missing counterfactual. The same inputs are never used to model a conventional small-cap target. Kocaarslan is explicit: "we do not claim that this sensitivity is unique to ESG-screened portfolios." He also writes that "distinguishing ESG-driven effects from broader size-related characteristics remains an important avenue for future research." The evidence therefore identifies credit-market sensitivity within one small-cap ESG index. Attributing that sensitivity to ESG screening requires a control the paper does not run.

Several implementation details remain unavailable. The tuned XGBoost settings include a learning rate of 0.01, depth of 6 and 1,000 trees, and the paper says cross-validation selected them. It gives no fold count, validation-window design or separate holdout period. The 80/20 alternative split is available only upon request.

We could not backtest the idea on our data. A faithful reconstruction requires the three provider-defined ESG index histories, or point-in-time constituent membership paired with the screening rules. We hold neither. Plain market-cap portfolios would remove the ESG mechanism being examined.

The paper is most useful as a risk-attribution note. Its two credit interactions suggest that an ESG label did not erase small-cap exposure to financing conditions during 2021-2024. The 0.99861 small-cap fit remains a same-day explanation, with no forecast embedded in it.