A one-standard-deviation rise in Zhang's earnings-call human-capital disruption score adds 0.55 percentage points of annualized idiosyncratic volatility within firm. The estimate has a t of 3.43 across 12,114 firm-years and 2,315 firms. Set beside the sample mean idiosyncratic volatility of 31.8%, the increase amounts to about 1.7% of the level. This belongs in a volatility model. Zhang's own twelve-month-ahead return coefficient is -0.0007 (t=-0.13).
We could not reproduce the score. Public materials omit the coding criteria, the 50 hand-classified excerpts that establish the economic boundary, the 729 training and 400 evaluation Claude Opus labels, and the fine-tuned DeBERTa ensemble. Nothing we ran can score a transcript as Zhang does. The backtest near the end uses a substitute measure and inverse residual-volatility sizing. It cannot test the paper's classifier or its reported coefficients.
Zhang's measure
The economic argument starts with human capital as a production input. A firm's execution becomes less certain when it struggles to hire the required skills, pays more to retain staff, absorbs unusual attrition, or works through restructuring. The expected effect lies in the firm-specific return distribution while market exposure stays unchanged. Zhang designs the tests around that split.
Human-capital disruption means a material disturbance specific to the reporting firm's workforce. The definition covers availability, skills, attrition and retention, wage and labor cost, safety and continuity, workforce restructuring, and consequential leadership transitions. Zhang excludes routine headcount disclosure, generic culture discussion, labor problems at another firm, and unsupported analyst questions.
Measurement has two stages. A workforce vocabulary identifies 1,835,450 transcript sentences. An embedding-plus-SVM screen set at 0.65 reduces them to 173,273 candidates from 36,159 calls, with precision 0.922 and recall 0.769. A three-model DeBERTa ensemble reviews each candidate alongside one sentence of context on either side. Training uses 729 Claude Opus labels. The acceptance threshold of 0.9208 was selected to deliver 90% precision on a separate 400-excerpt sample, where precision is 0.901 and recall is 0.585. Accepted excerpts are counted per 10,000 transcript words and then aggregated into Compustat fiscal years.
The source sample has 45,725 calls linked to CRSP and Compustat, covering October 2005 through May 2025 and drawn from two public transcript archives. Its annual panel spans FY2006 to FY2024, with 13,201 firm-years and 2,870 firms. The average score is 1.173 excerpts per 10,000 words, and 61.8% of firm-years have a positive score.
The annual within-firm estimates carry the most weight. Idiosyncratic volatility rises by 0.0055 (t=3.43), downside deviation by 0.0058 (t=3.92), and worst monthly return falls by -0.0046 (t=-3.75). Market beta changes by -0.0014 (t=-0.37).
At the call level, one standard deviation predicts roughly 0.50% higher log idiosyncratic volatility over trading days +2 to +43 after conditioning on matched pre-call risk. The t is 2.05 in the sample requiring that window to close before the next call. A score purged of every leadership and succession passage also predicts whether the incumbent CEO leaves before the next call. The increase is 0.391 percentage points from a 2.54% base rate (t=2.53), based on 715 exits across 27,691 intervals.
The return test lands at -0.0007, t=-0.13.
Next-year real outcomes are null as well. Employment growth has t=0.65, sales growth t=0.45, and ROA t=-1.14. Zhang gives the same reading in Section 6.6: "The current annual specifications therefore do not detect a common directional change in average real outcomes."
Separate from labor shortages?
Harford, He and Qiu, hereafter HHQ, previously published a FinBERT measure that counts labor-shortage sentences. The question is whether Zhang's wider workforce construct contributes anything beyond it. The distinction appears in the type of risk each measure tracks. Disruption loads on idiosyncratic volatility, while HHQ loads on beta.
The common annual sample contains 7,019 firm-years from 2006 to 2021. With the released HHQ measure and transcript-wide Loughran-McDonald negative and uncertainty frequencies included in the regression, disruption enters at 0.0065 (t=2.97). Removing every explicit shortage passage lifts it to 0.0068 (t=3.28). HHQ enters the same model at -0.0047 (t=-2.16). On its own, HHQ is -0.0014 with a t of -0.74, indistinguishable from zero.
Switch the outcome to market beta, using the same 7,019 observations, and the pattern reverses. Disruption is 0.0062 with a t of 1.30. HHQ is 0.0124 with a t of 2.30. Zhang treats the joint signs carefully, describing them as a decomposition of correlated text measures. His restrained conclusion is that the measures differ in their empirical relationships with systematic and firm-specific risk.
Their overlap has the same imbalance. Among 25,307 one-call firm-quarters, 62.2% of HHQ-positive observations are disruption-positive. Only 33.8% of disruption-positive observations are HHQ-positive, and the continuous correlation is 0.476.
Seven alternative classification rules probe the measurement choice. The grid changes the confidence threshold, requires agreement among independently trained classifiers, and removes shortage or leadership language. Accepted excerpt counts range from 30,294 to 78,003 around a baseline of 45,232, spanning -37% to +72%. Yet the idiosyncratic-volatility coefficient remains between 0.00503 and 0.00654, with a minimum absolute t of 3.16. Downside deviation ranges from 0.00561 to 0.00624, with minimum |t| of 3.73. Worst month runs from -0.00490 to -0.00439, with minimum |t| of 3.46. Firm-year rank correlations against the baseline range from 0.889 to 0.989.
Individual hard cases produce much weaker agreement. Across 120 boundary excerpts, Zhang and Opus agree 65.0% of the time. Cohen's kappa is 0.300, with an interval of 0.134 to 0.464. The grid matters because excerpt-level labeling noise appears to wash out before the firm-year ranking.
Timing weakens the trading case
The abstract acknowledges the timing problem and answers it through persistence. Risk is already elevated before the call, while the score predicts continuation of that firm-specific state over the next 42 trading days. Conditioning supplies the basis for the claim. Zhang writes that disruption predicts higher idiosyncratic volatility in each of the first two nonoverlapping 21-trading-day blocks after conditioning on the nearest pre-call risk realization and information available before the call.
The pre-call path fits that account. Combining the eight pre-call blocks into far, middle and near periods yields idiosyncratic-volatility coefficients of 0.63% (t=2.25), 0.69% (t=2.49), and 0.85% (t=3.36) as the call approaches. No discontinuity appears at the call date. Post-minus-pre changes have t-statistics of 0.59 for idiosyncratic volatility, -1.03 for downside deviation, and -0.36 for tail-loss magnitude.
Generic tone causes the short-horizon result to fade. Adding Loughran-McDonald negative and uncertainty frequencies at the call level lowers the 42-day coefficient from 0.00491 (t=2.09) to 0.00332 (t=1.59). Negative tone alone carries 0.02241 (t=7.36) in the annual model. Zhang reports both findings. In the 23,673 one-call HHQ quarters, he recovers a call-level estimate of 0.00653 (t=2.17). He also says plainly that the annual comparison supplies stronger evidence of incremental content. I share that ranking. The half of the paper that sounds tradable is the half I would decline to fund.
His own cross-sectional result imposes the harder limit. Section 6.3 states: "A cross-sectional specification with Fama-French 48 industry and year fixed effects, rather than firm fixed effects, yields an idiosyncratic-volatility coefficient close to zero. The annual result is concentrated in changes within a firm over time rather than in a stable cross-sectional ranking of firms." Zhang deserves credit for putting this in the body. It leaves the title and abstract with an unanswered question: why foreground a firm-specific risk measure when its only working form compares each firm with its own history?
A sort across 500 names should not be expected to produce a volatility spread from this score. The workable comparison is each firm against itself. The tests draw on a within-firm standard deviation of 1.343, versus 1.971 between firms. Subperiod estimates are 0.0094 (t=2.46) in FY2006 to FY2014 and 0.0033 (t=1.95) in FY2015 to FY2024. Two-way clustering reduces the headline t from 3.43 to 2.86.
Our substitute run
The public materials do not make the score reproducible. They leave out the coding criteria, the 50 hand-classified excerpts defining the economic boundary, the 729 training and 400 evaluation Opus labels, and the fine-tuned DeBERTa ensemble. We therefore could not score a single transcript in Zhang's manner.
Our run isolates the part of the design that can stand without text. It is a long-only book sized inversely to each stock's annualized Fama-French three-factor residual volatility. The estimation window covers the 42 trading days ending two days before rebalance. We used the 500 largest non-ADR US names and rebalanced monthly at the close. Costs were 10 bps one way plus four tenths of a cent a share. The period runs from 2020-01-01 to 2024-07-01. We omitted the disruption overlay that drops the top decile of forecast risk.
Over 2020-01-01 to 2024-07-01, our figures are total return 50.11%, Sharpe 0.56, Sortino 0.68, Calmar 0.23, volatility 19.89%, and maximum drawdown -40.76%. Read the volatility of 19.89% and the -40.76% drawdown first, since inverse-vol sizing makes a risk claim. The Sharpe of 0.56 comes from a window beginning with the March 2020 crash, monthly rebalancing, and one automated pass. It speaks only to our implementation.
Our 50.11% total return and 0.56 Sharpe have no comparable figure in Zhang's paper. The only return test I see there is the insignificant -0.0007 (t=-0.13) twelve-month-ahead coefficient, and I do not see any reported strategy performance. Because our run carries no text signal, it cannot test his claim or serve as a verdict on the authors' work.
The paper earns a place as a volatility-model term. Score each call, standardize the firm against its own recent history, and widen forecast risk by 0.55 percentage points of annualized idiosyncratic volatility for each one standard deviation increase in the score, lasting about one reporting interval. The 1.7% comparison uses the annual panel mean of 31.8%, rather than an individual stock's current level.
The missing test would change my view: a within-firm volatility forecast containing the score that sizes positions better than one excluding it. I do not see that test in the paper. Until somebody runs it, 0.55 percentage points against a 31.8% mean remains a small number on which to build a process.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.