A U.S. ESG ETF whose monthly returns correlate 1.00 with SPY is a fee decision wearing the costume of a portfolio decision. Chen prints that number in the full-sample correlation matrix, and it is the most useful thing in the paper.
Six funds and one Excel file
Chen compares six funds on monthly closing prices from December 2016 to May 2026. The window holds 114 monthly closes, so 113 monthly returns. Two ESG funds: ESGU (iShares ESG Aware MSCI USA) and ESGD (iShares ESG Aware MSCI EAFE). Two benchmarks paired to them by region: SPY for ESGU, VEA for ESGD. Then QQQ as a growth comparison and VICEX, the USA Mutuals Vice Fund, as a non-ESG comparison.
Chen does all of it in Excel, and all of it is descriptive. Monthly mean and standard deviation, annualized return and volatility, Sharpe using DGS3MO converted to a monthly rate. Cumulative growth of $100, and maximum drawdown as the worst peak-to-trough decline of that $100 series. A six-by-six correlation matrix, plus scatterplot trendline regressions of each ESG fund on its benchmark. Both of those are computed on the full period only. Five metrics get recomputed for each of four sub-periods: mean return, annualized return, annualized volatility, cumulative return and Sharpe. The windows are 2016-2019, 2020-2021, 2022-2023 and 2024-2026.
The paper is honest about where an ESG premium would have to come from. On one side, Friede and co-authors aggregate more than 2,000 empirical studies, and Whelan and co-authors aggregate 13 corporate meta-analyses covering 1,272 studies plus 2 investor meta-analyses covering 107 studies. Chen adds the cost-of-capital argument: heavy ESG inflows make these firms cheaper to finance. On the other side, the greenium. Demand pushes prices up and expected returns down. Hartzmark and Sussman find no evidence that high-sustainability funds beat low-sustainability ones. So the ESG trade is either a slow-burning risk premium or a valuation penalty already paid.
What Chen finds over the full sample: ESGU 1.06% a month, 11.99% annualized, 15.99% annualized volatility, Sharpe 0.60, max drawdown -26.40%. SPY 1.07%, 12.12%, 15.64%, Sharpe 0.62, max drawdown -24.80%. ESGD 0.59%, 5.95%, 15.51%, Sharpe 0.38, max drawdown -30.79%. VEA sits at 0.61%, 6.20%, 15.96%, Sharpe 0.39, max drawdown -30.70%. QQQ was the sample's best fund at 18.49% annualized, Sharpe 0.85, drawdown -33.07%. VICEX was the only loser: -2.09% annualized, Sharpe -0.26, drawdown -39.98%.
Chen reaches my spine himself. The abstract says their returns "appear largely explained by benchmark exposure, geography, and sector composition." Investors are also told not to expect "distinct, superior performance solely from ESG screening, as performance is largely based on other factors affecting the fund's actual holdings." I agree with that. My complaint sits elsewhere. ESGU is an ESG-screened version of an MSCI USA parent universe, so the near-identity with SPY is close to mechanical. And the paper reports no tracking error, no standard errors and no numeric beta to make it more than that.
The pairing is the contribution
Friede and co-authors aggregate corporate ESG-performance studies. Hartzmark and Sussman study fund flows and sustainability rankings. Chen's own step is the pairing discipline: each ESG fund gets measured against a benchmark in its own geography.
That matters for the comparison an investor actually makes. ESGU beat ESGD by 604 basis points annualized (11.99% versus 5.95%). Read alone, that looks like evidence that U.S. ESG screening works better than international ESG screening. SPY beat VEA by 592 basis points (12.12% versus 6.20%).
Almost the entire gap is region.
Chen backs this with holdings: ESGU's index puts over 30% into U.S. technology-prominent companies while ESGD sits heavy in Financials and Industrials, per BlackRock. Two funds with the same label, two very different absolute outcomes, and the label explains roughly a tenth of a percentage point of it.
Appendix C puts the full-sample ESGU-SPY correlation at 1.00 and ESGD-VEA at 0.99. Both funds are ESG-screened versions of MSCI parent universes, which is why the near-unit correlation is close to mechanical rather than a discovery. Chen writes that the comparisons indicated high R-squared values and that both beta values were close to 1. The two pair regressions are printed as Figures 4 and 5, with additional charts in Appendix B. No numerical alpha, beta or R-squared appears anywhere in the text. I wanted those values. No tracking error is reported, and no numeric beta either. So the 0.60 versus 0.62 Sharpe difference cannot be told from zero. The paper prints no standard error on any Sharpe or beta difference.
Which subperiod would you trade on?
The ESG edge changes sign across windows. In 2016-2019, ESGU ran a 1.03 Sharpe against SPY's 0.99, and ESGD 0.56 against VEA's 0.48. In 2020-2021, ESGU 1.25 against SPY 1.18. ESGD 0.47 and VEA 0.48 are the same fund for practical purposes. 2022-2023 was weak for all four paired funds: ESGU -0.15 against SPY -0.09, ESGD -0.20 against VEA -0.23. QQQ printed 0.04 in that window, VICEX -0.76. Then 2024-2026 flips the order: SPY 0.99 against ESGU 0.50, VEA 0.80 against ESGD 0.61.
Chen states the implication directly. Outperformance "was not consistent across funds, time windows, or metrics." He also flags that the windows were drawn around known regimes, which makes them an in-sample partition. The sample itself runs about 9.5 years, though the text calls it ten.
The 2024-2026 reversal is the one I would not build on. That Sharpe gap is a volatility gap. ESGU returned 14.91% annualized against SPY's 15.63% in that window. ESGU's volatility is printed at 20.64% against SPY's 11.10%. A fund whose full-sample monthly correlation with SPY rounds to 1.00 cannot run at 1.86x SPY's volatility. The paper reports no sub-period correlations. So the contradiction sits between the full-sample matrix in Appendix C and one row of Appendix D. It also sits against Chen's own sentence that both beta values were close to 1.
The 20.64% is the exact figure printed for ESGU in the 2022-2023 window as well. ESGD and VEA volatility repeat too: 11.14% and 11.45% appear in both 2016-2019 and 2020-2021. And in 2020-2021 ESGU is shown at 25.37% annualized with a cumulative return of 18.07% over 24 months. Cells in Appendix D do not reconcile. Chen's headline conclusion survives that, because the conclusion is that the pairs are near-identical. The specific claim that benchmarks have pulled ahead since 2024 does not.
One more measurement issue. Returns come from closing prices, and the paper never states whether they are dividend-adjusted. Within a pair the bias points the same way for both legs, so ESGU against SPY holds up. Across funds with different payout policies it does not. VICEX is the sample's only mutual fund, and a mutual fund's NAV drops on distributions. That matters for the fund the table ranks last on return (-2.09% annualized), worst on drawdown (-39.98%) and alone in negative Sharpe (-0.26), with the second-highest volatility at 17.30%. I would have expected a vice fund to hold up better than that.
We ran it as a switching book
We built a monthly switching version of Chen's pairing logic and backtested it from 2020-01-01 to 2024-07-01. Each month-end, for each pair, we form 12 monthly returns from real unadjusted closes. We also average the DGS3MO observations available in that month. Hold the ESG fund if its information ratio against the benchmark is positive. Hold it only if its Sharpe advantage is also positive. And only if its rolling beta sits between 0.90 and 1.10. Otherwise hold the benchmark. Half the book in the U.S. sleeve, half in the developed-international sleeve, executed at the next trading day's close. We charged 5 bps per unit of one-way turnover plus $0.004 a share with a $1 minimum, and modeled zero slippage.
The book returned 53.36% total over those 4.5 years. Sharpe 0.56, annualized volatility 21.10%, maximum drawdown -37.56%. Sortino 0.74, Calmar 0.27.
Our 0.56 is below the paper's 0.62 for SPY and 0.60 for ESGU. The setups differ: a switching two-sleeve book over 4.5 years against single funds over December 2016 to May 2026. I am not reading a winner out of that difference.
Two honest weaknesses in our version. The benchmark is the default holding whenever the three conditions do not all fire or history is short. Our sizing rules also conflicted: 50% sleeves against a 16.67% per-position cap. A 16.67% cap makes 50% sleeves infeasible, so the outcome depends on which one binds. We had no usable VICEX price history for 2020 through 2024, so none of the paper's VICEX statistics describe our run.
One automated pass is evidence about our implementation before it is evidence about the idea. The mechanism is where I would put the blame, and Chen's correlation matrix points at it. A 12-month information ratio computed on the difference between two series whose full-sample correlation is 1.00 is an estimate on almost pure noise. Every switch pays 5 bps of one-way turnover, with no slippage modeled on top.
A total-return series with a printed tracking error, and standard errors on the ESGU-minus-SPY difference, would change how I read this. If the 9.5 point volatility gap in the 2024-2026 row is real, ESGU deserves a fresh look. That gap and a full-sample correlation of 1.00 with SPY describe different funds. I do not believe the cell.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.