After the authors' own trading costs, the sentiment-driven Black-Litterman book trails a market-cap-weighted basket of the same eight stocks on Sharpe, 1.55 against 1.73. The cover's 106.58% annualized return is gross. It comes from annualizing 117 trading days that include China's late-September 2024 stimulus rally.
We have no SSE 50 price history or Eastmoney Stock Bar feed carrying view and reply counts. Nothing below is a reproduction. Replacing those inputs with US equities and US news sentiment would create a different study.
How the book is built
Sun, Fu and Liu collect posts from Eastmoney Stock Bar, a retail forum for discussion of individual Chinese stocks. Their universe contains eight SSE 50 constituents across real estate, banking, pharma, energy and semiconductors. The final sample has 387,953 comments from July 1, 2022 to December 31, 2024, after the removal of 4,725 official news posts and 2,508 duplicates or blanks.
They score every comment two ways. The dictionary method uses a combined Chinese financial lexicon containing 9,661 words, split into 3,551 positive and 6,110 negative terms. Its score is positive minus negative over total. The second method is BERT-base-Chinese, fine-tuned on hand-labelled titles from Eastmoney SSE 50 ETF posts: 3,170 negative, 3,024 neutral and 4,452 positive. Each daily index weights a post by its view count plus its reply count, giving louder posts more influence.
An LSTM takes the sentiment index alongside OHLC and volume, then forecasts the next closing price for each stock. Hyperparameters are grid-searched stock by stock. Black-Litterman combines returns implied by market-cap weights with the investor's stated views. In this implementation, the views are log returns derived from the LSTM forecasts and entered as absolute views Q. P is the 8x8 identity, tau is 0.025, and Omega equals tau times P Sigma P transpose. Sigma comes from a rolling 200-day covariance of excess returns. The risk-free rate is 2.39%, matching the 2024 average yield on the 10-year Chinese government bond. Starting from one yuan, the strategy rebalances daily across 117 trading days from July 11 to December 31, 2024.
The proposed mechanism makes sense. Retail chatter may reach prices with a lag that a language model can detect while a price-only model cannot. The harder question is whether 117 days reveal that mechanism.
A return built around the rally
Before costs, BERT_BL produces a 106.58% annualized return, Sharpe 2.64, Sortino 4.18 and max drawdown 6.84%. Dict_BL records 38.66% with Sharpe 1.21. The no-sentiment version returns 28.60% at Sharpe 0.94.
The benchmarks cover the identical 117 days. Market-cap weighting returns 43.28% with Sharpe 1.74, while equal weights deliver 40.41% at 1.40. The SSE 50 returns 28.51% at 1.09. Every passive benchmark annualizes between 28% and 43%. Mean-variance, the fourth comparator and an active strategy, earns 21.40% gross with Sharpe 0.83. Its 11.66% drawdown is the worst in the table.
No normal half-year in Chinese equities produces passive returns like these.
The authors give a direct explanation: positive stimulus arrived in late September 2024, sentiment rose, and the market rallied before a broad correction. Their Pagan-Sossounov regime classification marks a bull phase from September 13 through October 8, 2024. It lasts eleven trading days. A 35-day bear phase follows and runs to November 25, leaving 71 of the 117 days outside either identified regime.
None of the portfolio return figures comes with t-statistics or standard errors. The 2.64 Sharpe covers the full period on a gross basis. An eleven-day bull phase inside that figure is far too brief to support the claim.
Costs take away more than half
The cost assumptions deserve credit. The authors charge 0.03% commission on both sides plus 0.1% stamp duty on sells, matching the real A-share schedule. They also state plainly that all strategies begin from the zero-cost assumption.
BERT_BL falls from 106.58% to 51.87%. Its Sharpe drops from 2.64 to 1.55, and drawdown rises from 6.84% to 8.23%. Dict_BL nets 4.69% with Sharpe 0.22, far below the SSE 50's 28.51%. No_Sentiment_BL becomes negative at minus 4.48% annualized and Sharpe minus 0.10. Market-cap weighting barely changes, moving from 43.28% to 43.02%, because its weights drift slowly.
Once costs are charged, passive market-cap weighting wins on risk-adjusted performance. Its Sharpe is 1.73 against 1.55 for BERT_BL, with Sortino 2.41 against 2.47. BERT_BL still leads on raw return, 51.87% versus 43.02%, while losing heavily on drawdown, 8.23% versus 0.75%. Removing sentiment leaves the Black-Litterman machinery with a negative net Sharpe. The apparent edge therefore lies in an LSTM forecast fitted over one five-and-a-half-month window that contains an eleven-day policy rally.
The 0.75% drawdown needs separate scrutiny. The authors credit large-cap stability for lower portfolio volatility. Yet a large-cap tilt does not explain a 0.75% peak-to-trough loss across eight names during a period when the SSE 50 itself drew down 9.50%.
How much signal is there?
Fine-tuned BERT scores 74.69% accuracy on its labelled ETF-title test set. Accuracy drops to 52.3% on 1,000 randomly selected comments from the scraped corpus, labelled by hand. The authors say most errors confuse polarity with neutral sentiment instead of flipping positive and negative. Their confusion matrix nevertheless assigns 100 of 438 actual-negative comments to the positive class, or 22.8%. The claim that negative-to-positive misclassification is extremely low therefore fails on their own figures. The LSTM receives sentiment from a classifier that is correct about half the time on the text where it is used.
Same-stock Spearman correlations between sentiment and log returns reach a maximum of 0.2016, achieved by CSCEC under the dictionary index. BERT sentiment is significant at the 1% level for five of the eight names: China Unicom at 0.1668, Hengrui at 0.1107, Yili at 0.1231, CSCEC at 0.1633, and AMEC, the semiconductor equipment maker, at 0.1190. Results for Sinopec and Poly are negative and insignificant, at minus 0.0519 and minus 0.0385. Correlations from 0.11 to 0.17 across five of eight names offer an unconvincing explanation for a 4.18 Sortino.
The MSE evidence is cleaner. It is also the paper's most interesting result. Adding sentiment reduces LSTM test-set error for all eight stocks. CSCEC improves from 0.2777 without sentiment to 0.0316 with BERT. Sinopec moves from 0.0766 to 0.0185. BERT outperforms the dictionary on six of eight stocks, while the dictionary leads for Sinopec, 0.0151 versus 0.0185, and China Unicom, 0.0119 versus 0.0277. These forecasts target price levels in a trending series, so lower MSE can occur without directional edge. The paper reports no hit rate anywhere.
Tau barely matters
Moving tau, the weight assigned to views, from 0.001 to 1.0 leaves the net annualized return in a narrow range of 51.27% to 52.00%. Sharpe stays between 1.52 and 1.55. This diagnostic shows the views overwhelming the prior across three orders of magnitude. Shrinkage toward market equilibrium, Black-Litterman's central function, is not binding in this setup. Describing the strategy as an LSTM forecast book sized through covariance would lose little.
What evidence would settle it?
The abstract says the model survives transaction costs, parameter changes and different market environments. Its regime evidence consists of an 11-day bull window and a 35-day bear window, presented only through a cumulative-return chart. Neither subperiod has a reported return, Sharpe or drawdown figure. The strongest case is relative: within the same window, BERT_BL outperforms the other BL variants and the benchmarks because BERT responds faster to sudden sentiment changes. Performance through one policy shock carries some information. Across 46 classified days without a dispersion estimate, luck remains impossible to separate from skill.
A walk-forward study over the full 2022 to 2024 sample would answer the question. It would exclude the stimulus window and report directional hit rate beside the MSE results. If BERT_BL still beats the market-cap benchmark after paying 0.03% on buys and 0.13% on sells, a 0.16% round trip, the mechanism has earned the claim.
The narrower contribution survives. Forum sentiment measurably improves LSTM price-level forecasts for eight Chinese large caps, and the gain is large enough to move a Black-Litterman book from a negative net Sharpe to 1.55. The result is highly sensitive to costs. Dict_BL nets 4.69% against the index's 28.51%, while the no-sentiment version turns negative. All of that is publishable. The annualized return on the cover does not supply the evidence.