On six S&P 500 pairs, headline sentiment cost the strategy return. The best of the authors' sentiment-aware pairs trading strategies averaged a $69.77 return per $100 position across all pairs; plain Bollinger bands averaged $86.74. Lee, Bernard, van Heeswijk, Coita and Machado acknowledge the gap. They make their case on risk instead, where the evidence is less convincing.
The trades behind the comparison
The idea follows familiar behavioural pairs logic. When two stocks that usually move together diverge, news tone might distinguish an overreaction from a lasting change. The pairs are PEP-KO, HPQ-DELL, UAL-AAL, MA-V, CVX-XOM and MLM-VMC. Selection required an Engle-Granger p of 0.05 or less, or a three-month correlation above 0.8. Five pairs meet it. UAL-AAL (p 0.5737) and CVX-XOM (p 0.6604) qualify on correlation alone (0.8582, 0.8532); PEP-KO qualifies on cointegration alone (correlation 0.6675). HPQ-DELL failed both (0.1600, 0.5698) and remains as a control.
Prices are one-minute adjusted closes from LSEG for one year starting 13-03-2025. Google News headlines were found by company name and scored using VADER's compound score. One version retains every score; the other removes neutral scores (magnitude below 0.1). Each stock had 147 to 469 headlines over the year, falling to 75 to 356 after neutral scores were removed.
The benchmark trades the log spread against bands. Its window is three times the spread's half-life, estimated on the training 80%, while a rolling OLS supplies the hedge ratio. Sentiment enters through three variations: "Weighted" adjusts band width using the in-sample correlation between sentiment and volatility; "Semi-Predicted" uses regression forecasts of each leg's variance to set that width; "Fully Predicted" forecasts the mean and covariance as well. Separately, an LSTM forecasts the spread's two-minute maximum and minimum, then trades against those bounds. The assumed costs are 0.5% for commission, 0.5% for bid-ask and 2% a year to borrow. The test lasts 68 days.
Does the risk claim hold pair by pair?
The authors say Fully Predicted "generated consistently higher Sortino ratios, and lower maximum draw-downs." Figure 7 supports that wording with averages across pair-model combinations. The printed pair rows give a less consistent picture.
MLM-VMC provides a direct check because its row uses average sentiment with neutral headlines dropped. Sortino moved from 0.292 to 0.288. Drawdown fell from 0.017 to 0.002, while trades fell from 172 to 11 and return from 58.99% to 24.47%. MA-V favours the authors' claim: Fully Predicted with average sentiment and neutral headlines dropped recorded a Sortino of 1.029 against 0.608, and drawdown of 0.016 against 0.025. The sentiment-variance version split the outcomes. On DELL-HPQ, Sortino rose from 0.025 to 0.492, alongside a drawdown increase from 0.023 to 0.050. PEP-KO went the authors' way on both measures, with Sortino 0.082 against 0.018 and drawdown 0.024 against 0.029.
Eleven trades against 172 can reduce drawdown without a better signal. A comparison holding trade count or time-in-market fixed would help separate the two. We did not find one.
Selection matters here too. The table's "best" sentiment model for each pair is chosen by return and by Sortino across thresholds from t=0.1 to t=5, both sentiment versions, and means against variances. Those choices are evaluated on the same 68-day test window. Some Sortino winners barely trade: 3 trades for PEP-KO and 2 for CVX-XOM.
Random inputs kept pace
The authors deserve credit for running a random normal input (standard deviation 0.34) as a control. It performed well. In their words, "the highest overall return $314.52 was achieved by a strategy that utilised random sentiment highlighting the difficulty of determining the influence of sentiment data from chance" for trading applications. The table gives that row as 314.36%. It belongs to MA-V, where Fully Predicted on noise beat the 301.52% result for real sentiment.
The authors answer that a sentiment-aware strategy delivered the highest return on PEP-KO and DELL-HPQ. Their paper calls that one third of the pairs. It also counts MA-V as a third sentiment win over the benchmark, though noise beat every real-sentiment variant there. DELL-HPQ's return winner, Weighted at t=0.1, earned 50.74% against 27.98%; its drawdown was 0.127 against 0.023.
The LSTM adds little support, and the authors reject both LSTM hypotheses themselves. Noise produced the lowest RMSE on 3 of 6 pairs and the highest average LSTM trading profit (52.264, Sharpe 0.0940). No LSTM variant beat $86.74. On MA-V, RMSE reached 1503.94% of the average spread. The conclusion calls the sentiment LSTM's lower accuracy "improvements," which reads like a drafting slip.
The regression results pose a more direct problem. The text says multivariate analysis "confirms that sentiment has predictive power for all stock pairs." Yet Tables 11 and 12 say "Accept null hypothesis" for all four sub-hypotheses, for every pair shown under both sentiment versions. Univariate fits reach 0.9978 (XOM), and pair-level fits reach 0.9961 (PEP-KO). Both emerge from searching 729 time combinations with 100 observations per regression and a p < 0.10 screen. Across the univariate search, average adjusted R² peaks with 28-day sentiment windows and 6 to 24 hour lag and price sampling. If, as we read the setup, the 100 points follow the price interval, neighbouring sentiment averages reuse nearly all their headlines. The conclusion describes the findings as "sample-dependent associations rather than as definitive causal or universally stable predictive relationships," and treats the 2 to 4 day peak as exploratory. That wording fits the evidence better than the abstract's.
The MA-V cost question
MA-V's benchmark returned 268.97% in 68 days across 233 trades, with average hold rounded to 0.0 days. At the stated 0.5% commission plus 0.5% spread, even one charge per trade adds up to about 233 points before compounding. An intraday round trip between two payment networks would therefore need to gross about 1.5% before costs. We could not reconstruct that from the described method. The paper leaves unclear how often it applies costs and whether its drawdowns are fractions or percentages.
PEP-KO needs separate caution. Its sentiment result more than doubled the benchmark, 41.50% against 19.44%. We have also watched that spread stop mean-reverting out of sample before (/articles/from-cointegration-to-out-of-sample-failure-a-pairs-trading-case-study-on-pep-ko).
Before this reaches a book
A commercial news-feed rebuild would begin with sparse input: 147 to 469 headlines per stock per year, at most about one a day, supplying minute-level features. Company matching comes before tone. The authors concede that searching by company name may have missed relevant news. Ambiguous names such as 'Visa' and 'United Airlines' in their Table 2 could also collect irrelevant headlines; HP Inc was searched as 'Hewlett Packard'.
The Fully Predicted variant should then be frozen on the training 80%, run once at the benchmark's trade count, and compared with a distribution of random-input runs. If MA-V's Sortino gain, 1.029 against 0.608, survives, sentiment earns a place in the band width. For now, the authors' own average favours the news-blind benchmark at $86.74 per $100, with that figure leaning heavily on the 268.97% MA-V row whose costs remain unexplained.