Cleaner news topics have yet to earn a trading edge. At 60 topics, the sentence transformer reaches mean NPMI of 0.3469 against LDA's 0.2875. NPMI scores how often a topic's top ten terms occur together. The paper's abstract concedes that its tests do not establish outperformance.
LDA treats an article as word counts, without word order. The paper instead uses a frozen sentence transformer to encode running text, then clusters the passage embeddings into topics. The trading question is whether that context-aware text layer prices stocks better when both layers enter the same factor model.
From news attention to stock weights
Foley, Hartadi, Prakash and Vatsal take the narrative asset-pricing pipeline of Bybee et al. (2023) and change only the text layer in their main comparison. Each article receives weights across L=60 topics. Averaging those weights produces a daily attention series for each topic; today's attention minus its trailing 5-day mean is the narrative shock.
For each stock, the model uses a 252-day lookback and a 60-day half-life to estimate exponentially weighted covariances between daily returns and the 60 shocks. The 60 covariances, plus an intercept, become stock characteristics. Sparse IPCA maps them to loadings on K=3 latent factors. Its group penalty can remove an entire narrative from this instrumented factor model. Each month, the tangency portfolio of the three factors is translated back into stock weights.
The proposed premium comes from priced state variables. In the authors' example, an oil-supply shock affects producers and airlines differently. Stocks that comove with the shock carry exposure to it; a priced exposure would pay the factor portfolio.
The news sample contains 394,661 FNSPID articles dated January 2015 to December 2023. The CRSP panel has 95,071 stock-months from August 2015 to January 2024. Estimation begins with the first 61 months. The remaining 41 months, September 2020 to January 2024, are out of sample, with expanding-window refits every 12 months.
The inputs differ alongside the models. LDA uses Gibbs sampling over a 15,000-term vocabulary and the full cleaned article, capped at 4,000 characters. The transformer branch feeds at most three lead chunks, each cut to about 110 words, into frozen paraphrase-MiniLM-L3-v2. Mini-batch k-means groups the embeddings into 60 clusters. A softmax at temperature 0.07 gives passages their topic weights.
The transformer raises mean NPMI from 0.2875 to 0.3469. Its OOS total-return Sharpe is 0.3059, against -0.0810 for LDA. Two subsequent extensions lift an excess-return Sharpe first to 0.6605 and then to 1.0334.
What did the coherence gain measure?
The 0.0594 NPMI gap includes two differences the authors themselves flag. The transformer reads lead passages while LDA reads the cleaned article, and the branches use different rules to rank topic terms. LDA uses fitted term probabilities; the transformer uses class-based TF-IDF. A topic-level Welch test of the NPMI gap gives p=0.042. The authors also caution that topics drawn from one corpus need not be independent.
The ranking rule can move the score substantially. In the spherical transformer run, changing only the term head from full-soft to sparse-soft takes NPMI from 0.3702 to 0.4691. That 0.0989 gain exceeds the entire 0.0594 LDA-versus-transformer gap. The paper does not measure how much of the latter gap comes from ranking. Yet daily attention under the two heads correlates at a minimum of 1 minus 6×10^-14. The authors call the NPMI increase "produced entirely by the ranking of terms", which fits those near-identical attention series.
The market comparison
The equal-weight market comes out ahead.
Over the same 41 months, an equal-weight portfolio of the same universe records a total-return Sharpe of 0.8166. The single-horizon narrative portfolios reach -0.0810 with LDA and 0.3059 with the transformer. The difference between those two branches is 0.3869 Sharpe, while a paired test of mean returns gives p=0.22. The authors estimate the standard error of an annualized Sharpe at roughly 0.55 over 41 months. Their measured gap is smaller than one standard error.
The appendix shows how readily the ordering changes. At L=80, single-horizon LDA scores 0.7675 against 0.7219 for the spherical transformer; these L=80 runs use spherical clustering. Adding 20 topics moves the LDA baseline from -0.0810 to 0.7675. The authors judge the two layers indistinguishable on the single-horizon specification at L=80. The same reading applies at L=60.
How the 1.03 Sharpe arose
Moving from 0.3059 to 0.6605 changes three choices together: covariance half-lives of 20, 60 and 120 days (181 instruments), total returns replaced by excess returns, and a new penalty grid. Switching to spherical k-means then contributes 0.3729. The paired t for that last change is 0.692 (p=0.493); Newey-West gives p=0.521. A block-bootstrap 95% interval for the Sharpe difference runs from [-0.903, 1.682].
Before trading costs, the raw spherical portfolio has 81.68% annualized volatility and a 62.23% maximum drawdown. The authors disclose that they chose the extension after inspecting the same 41 OOS months. Instrument scaling uses the full panel, including OOS months, and the text models were fit on the full 2015 to 2023 corpus. In the spherical runs, the group penalty keeps all 60 narratives at every refit.
At L=80, the headline specification returns a Sharpe of 0.7156. The authors note that it remains above multi-horizon LDA's 0.6125, with a gap statistically indistinguishable from zero. It also nearly matches the equal-weight market's excess-return Sharpe of 0.7164.
The paper says "the available tests do not establish outperformance" while seeing "some promise" in context-aware text. Its request for stricter tests and broader datasets is right. On Sharpe, the single-horizon ranking flips as L moves from 60 to 80. The coherence lead persists at every L the authors tried.
A chronological rebuild
The authors' future-work section calls for fitting vocabulary and topics on pre-September 2020 text and scaling instruments on the training sample. It also calls for identical passages and term-ranking rules across branches, with excess-return portfolios reported across topic counts and seeds. A chronological rebuild would need to beat the equal-weight excess Sharpe of 0.7164 on the same window, with its text layer frozen before September 2020.
Until that happens, the transformer's demonstrated gain remains in the labels.