The 2.85 net Sharpe is unconvincing without a turnover figure. Kirtac has paired a strong evaluation protocol with a portfolio result that the protocol, as published, cannot support. The framework should be adopted.
MFAST, the author's name for his Market-Friction-Aware Sentiment-to-Trading pipeline, turns timestamped news into a portfolio through an ordered sequence of operators. Each article is assigned to a tradable firm. Repeated wire copy is screened out, a language model scores the text, and its probability is calibrated on a validation set. The pipeline then ranks firms within each day's eligible universe and builds a value-weighted quintile long-short book. Costs and a participation cap come before the Sharpe ratio. Grossman-Stiglitz supplies the economic argument: public news is free to read and expensive to interpret, allowing a model that processes negation and forward-looking guidance faster than the marginal investor to collect the processing rent, especially where frictions delay incorporation.
The paper merges Refinitiv News Analytics with CRSP daily equities from January 2010 to January 2026. Its funnel begins with 3,129,924 raw items, falls to 1,985,135 single-firm stories, then to 1,122,475 after a five-day cosine novelty screen at 0.80. The final sample has 973,481 tradable items across 3,452 firms. Tradability requires positive quotes, 1,000-share daily volume, $50,000 daily dollar volume and quoted spreads under 20%. Training covers 2010 to 2023. January to May 2024 serves as a validation and release buffer. The reported out-of-sample period runs from June 2024 to January 2026 and contains 190,236 articles. It begins after both the disclosed LLaMA-3 family data-freshness cutoff and the checkpoint's public release.
We cannot reproduce any of this. Refinitiv News Analytics and CRSP are unavailable to us, we have no historical bid/ask quotes, Kyle lambda cannot be reconstructed from OHLCV bars, and our news tables begin around 2020.
Labels use the sign of the three-day cumulative excess return over the execution-aligned window, measured from execution day through two days later against the CRSP value-weighted market. Six scorers use identical labels: LLaMA-3-8B and OPT-1.3B adapted with LoRA, RoBERTa-base, BERT-base and FinBERT fully fine-tuned, plus the Loughran-McDonald lexicon.
LLaMA-3 leads every reported signal-quality measure. Accuracy is 0.787, AUC is 0.846, Brier is 0.151 and expected calibration error is 0.032. In next-day fixed-effect regressions, its coefficient is 0.312, with a t of 6.44 and a within R2 of 0.052 across 190,236 observations. The dictionary records 0.049 with t of 1.31. Operating cost goes the other way: LLaMA-3 processes 8.9 articles a second on 4 x A100, compared with 117.6 for FinBERT.
The case for MFAST
Kirtac's main contribution is the order of evaluation. MFAST forces the questions that follow the F1 score, and the design matrix states plainly which test addresses each requirement. Timestamps are converted to Eastern Time and assigned to the next feasible decision point. An article released at 10:15 therefore cannot claim that day's return. A boundary duplicate audit also removes same-firm near-copies within 20 trading days, preventing the model from training on an event and later being scored on its rewrite.
The feature ablation is the paper's most revealing table. It also pushes against the fashion for multimodal fusion. Text alone produces 0.846 AUC, while price alone reaches 0.565, liquidity 0.548 and metadata 0.552. Combining everything raises AUC to 0.866. All the structured data in the pipeline therefore adds 0.020 of AUC.
Can the 0.34% daily return survive turnover?
The headline portfolio is a daily-rebalanced, value-weighted quintile long-short book covering a universe of up to 3,018 firms. Results are net of 5 basis points one-way, with trading capped at 10% of daily dollar volume. The book earns 0.34% mean daily on 1.89% vol, suffers a drawdown of 12.3%, and compounds to 180% over roughly twenty months. A block bootstrap places the Sharpe CI at [2.31, 3.34]. Daily alpha is 0.182% with CI [0.109%, 0.252%], while the deflated Sharpe is 2.19 with p of 0.014.
Turnover appears among the implementability tests, yet the paper supplies no figure. That omission controls the reading of the entire portfolio result. A daily quintile rebalance across a three-thousand-name news universe creates high turnover by construction, and net return equals gross return minus c times the sum of absolute weight changes. The 5bp assumption cannot be assessed without that sum. The paper says the cost ranking remains stable from 0 to 50 basis points, but gives the finding in prose rather than as a grid.
Factor alphas, subperiod panels, capacity diagnostics and the liquidity gradient receive the same treatment. The liquidity heterogeneity section says coefficients increase from high-liquidity to low-liquidity stocks, with the steepest slopes for LLaMA-3 and OPT. It reports no coefficients.
The limitations section acknowledges that the execution model omits intraday liquidity, queue position, hidden liquidity and strategic interaction because it is not a full order-book simulator. In the same passage, the paper calls the model realistic. Realistic against what turnover?
Value weighting is presented as the defence against microcap contamination. Yet the capacity section concedes that the 10% participation constraint binds most often among smaller and less liquid names, precisely where the mechanism test locates the edge. Neither side receives a number.
Short borrow carries no cost. Long leg 1.72, short leg 1.48, so roughly half the book relies on it.
A benchmark at chance
Loughran-McDonald classifies the sign of the three-day excess return with 0.503 accuracy and 0.512 AUC. A coin. Gains against a baseline at chance are difficult to interpret.
The error analysis uses 1,200 stratified test articles, balanced across agreement and disagreement cases, and relies heavily on that comparison. LLaMA-3 exceeds the dictionary by 0.365 on contrastive clauses and 0.335 on forward-looking guidance, versus 0.171 on simple polarity. The pattern is credible and provides the paper's strongest mechanism evidence. Its level still depends on a benchmark unable to distinguish a positive three-day return from a negative one.
FinBERT offers the cleaner comparison. Across the same 1,200 articles, the compositional gaps shrink to 0.109 on contrastive clauses and 0.114 on guidance.
Contamination remains after the safeguards
The paper handles this issue carefully and states the limit directly: the safeguards "do not prove the absence of all document-level pretraining exposure." It uses static checkpoints, disables retrieval and browsing, and starts the test window after the disclosed freshness cutoff and public release. Those are appropriate defences.
Selection on pre-test data remains. Hyperparameters, probability calibration, aggregation rules, transaction-cost levels, portfolio cutoffs, holding periods and participation caps all come from the training and validation periods. The strategy family also follows the author's earlier work with Germano. The deflated Sharpe adjusts for search within this paper. Search across a research programme remains outside that correction.
The public replication arm supplies the piece's most useful disclosure. GDELT headlines matched with Yahoo and Stooq prices achieve 0.612 accuracy and 0.651 AUC. The open friction-aware portfolio records a Sharpe of 0.92 on 41% cumulative. Kirtac explicitly describes the open data as noisier in timestamps and entity identifiers, and says this arm validates the pipeline rather than the proprietary result. Fair. The distance between 0.92 on GDELT headlines and public prices and 2.85 on Refinitiv and CRSP is still wide. The author attributes it to those noisier public timestamps and identifiers.
The primary study remains beyond reproduction
Refinitiv News Analytics and CRSP are unavailable to us. Without historical equity bid/ask quotes, the quoted-spread eligibility screen and spread-based cost estimates must be replaced by bar-based liquidity filters and an assumed cost schedule. Kyle lambda cannot be reconstructed from OHLCV, so price impact has to be approximated using dollar volume, volatility and a participation assumption. Our news tables begin around 2020, which prevents us from matching the 2010-2026 sample or the post-freshness-cutoff design.
We are running a version on our own data. It is not a test of the paper's numbers.
The framework is what traders should take from this paper. Kirtac's sequence, observability before calibration before ranking before costs before capacity before inference, is the evaluation order news-sentiment research should have used for a decade. One disclosure would change our view of the portfolio result: two-way turnover, together with the associated cost drag, for the LLaMA-3 book over the twenty-month test. Publishing it would make the 2.85 checkable. Leaving it absent means the framework audits everything except the part that determines the answer.