The echo filter is the sign rule a desk could test now. Give it a timestamped news tape with embeddings, and it can identify repeated coverage to fade. The paper's other trade, fading an institution when its words conflict with its position, depends on data most desks lack.
Why the sign reverses
Alzahrani models an informed institution that speaks about an asset's value and trades it before a partly credulous crowd reacts. Three parameters govern the choice. Credulity, φ, measures the crowd's price response per unit of message. The institution pays λ in price impact on its trade. A misstatement costs k/2 times the squared gap between the message and true value.
The linear-quadratic solution divides into regimes. If φ² < 2λk < φ, the institution issues what the paper calls a false alarm: it talks the asset down and buys. Once credulity exceeds one (and 2λk > φ², so an optimum exists), exaggeration takes over. The institution talks the asset up and sells.
The useful result lies beneath those regimes. If the trade is optimal given the statement, the forward return is λx − φε pathwise. Here ε is the noise in how the public hears the message. For any distribution of value, Cov(return, Say) = λ·Cov(Do, Say) − φσ²_ε. Say denotes the institution's statement; Do denotes its revealed position. Negative Say-Do covariance makes the words predict returns with the wrong sign. Price and volume cannot reveal that covariance: neither identifies the buyer.
The paper also separates fresh news from echo, meaning articles that repeat earlier ones. It models the articles as a Hawkes cascade and computes an exact posterior for each article's parent. Arrival times and a von Mises-Fisher similarity on embeddings enter that calculation, which uses only articles already published. In the theory, a pure garbling of the other channels gets zero Bayes weight. The crowd's weight on echo should be reversed.
The tests are simulations. The main market contains 300 firms and 60 events per firm, with 30 training and 30 test events each. Seeds and variants make up 29 distinct simulated markets. A quarter of institutions are non-strategic; the others divide evenly among shading, false alarm and exaggeration. Half of all institutions consequently have words that should be faded.
Can a news tape replace the missing trade?
For echo, it gets close. Declustering estimates an echo share of 0.422 versus the true 0.427, and concentration of 30.27 versus the true 30. A timing-only estimate gives 0.368 against 0.430 because it mistakes echoes for news.
Echo sentiment predicts returns with a significantly negative sign in all 29 markets. The firm-clustered t in the main market is −8.0. Across the three robustness markets, timing alone weakens the echo t from −7.5 to −6.0. Positions are unnecessary to run the filter itself, although the paper controls for Say and Do in its echo regression.
For the institutional rule, the trade record remains indispensable. Rolling Say-Do correlation identifies false-alarm events with an AUC of 0.895. As a switch between fading and following a firm's words, it is right 85% of the time after 5 past events and 96% after 40. Those results all rely on a Do series attributable to the speaker.
Even within the simulator, the reversed-words result has a weak point. Say predicts returns at t = −3.3 in the flagged group. Control for echo sentiment and the t moves to +0.2. Alzahrani reports this himself, interpreting the result as the statement's price impact passing through its repetitions. With a saturating crowd, the controlled test takes the wrong sign in all three seeds. Do's coefficient reaches only t = +0.9, partly because Say and Do have a correlation of −0.87 among false-alarm firms.
Echo holds in 29 markets; added features gain 0.002 of IC
A linear model using raw Say, Do, tone, article count and price reaction produces an out-of-sample IC of 0.169. Add all the theory features and it reaches 0.171; no seed gains more than about 0.005. Boosting goes from 0.109 to 0.118 in the main market, with gains above 0.001 in 18 of 29 markets. Fading price reaction alone yields 0.162.
The unit-weight theory composite reaches 0.049. Its absorption term, the price move net of the narrative-implied move, has the wrong sign in this market. Drop that term and the composite doubles to 0.096.
Alzahrani is measured about the comparison. True mispricing gives the oracle an IC of 0.181, leaving about 0.01 above the linear baseline. The paper says the comparison has little power and calls "its signs and the diagnostics" its clearest contributions. By that standard, the theory contributes little to prediction here. Theorem 1 makes the optimal forecast linear in the voices; simulated price is only approximately linear. Even so, a free regression finishes within about 0.01 of the oracle. The theory's contribution is an expected sign for each weight and a warning when the fitted weight goes the other way.
The calibration matters. In a pilot run, fading tone alone achieved an IC of 0.89, while prices moved about 3.3 times as much as fundamentals. The author reduced per-article news impact from 0.12 to 0.04. He then chose late noise to bring the linear baseline near 0.17, a level the paper calls far more predictable than any real market. The intended comparison is between methods. We found no Sharpe, cost or turnover figures in the reported tables, although Sharpe appears among the metrics.
What a live test would require
We could not backtest the Say-Do rule. We lack institution-level holdings or flow and a feed of analyst notes and rating changes. The rule needs positions tied to the institution that made each statement; OHLCV cannot say who bought.
The echo half could run on a news archive with embeddings.
Testing the central rule point in time would require each statement's issuer and timestamp, plus that issuer's position series. The paper names holdings, insider trades, futures positioning and order flow as possible sources. The disclosure delay matters too. In the simulator, Do arrives with the event as the trade plus noise with standard deviation 0.5. We did not find a disclosure lag modelled.
In the closest variant, only 30% of the trade is disclosed and the noise doubles. Detection AUC falls from 0.941 to 0.857 across the same three seeds. A lagged holdings report would arrive after the echo cascade it is supposed to help discount. The 85% switch accuracy also requires five matched Say-Do pairs for an issuer before the first call.
The paper proposes testing echo and Say-Do predictions first. A desk could run the echo test: its inputs exist outside the simulator, and its result was significant in all 29 markets (news was significant in 21 of 29).