The last hourly USD/CAD price beats every level forecast in this central-bank news study. Kodeih, Alnaggar and Cevik have a useful feature result, but their case for directional and probabilistic forecasting is thinner.
Why news timing might matter
Fed and Bank of Canada communication can shift rate expectations, then USD/CAD. The authors ask whether hourly text features capture policy news before the exchange rate fully reflects it.
They gathered articles from TheNewsAPI and RSS feeds, filtered them using monetary-policy terms, then deduplicated them. The result was 927 articles from 2024-06-13 to 2026-06-12, matched to 17,517 hourly USD/CAD observations from Yahoo Finance. GPT-4o-mini assigned each article a hawkish, dovish, neutral or mixed label for the Fed and the BoC separately. Those labels and their timestamps produced a Fed-minus-BoC sentiment differential, impulse and exponential-decay variants, rolling averages, per-bank article counts, hours-since-news, and is_news_hour, which marks an hour when policy news arrived. USD/CAD lags, intraday volatility, the US-Canada rate spread and WTI supplied the market inputs.
Extra Trees, XGBoost and LightGBM forecast the level directly at 1, 3, 8 and 24 hours. Rolling windows held roughly 4,324 training hours and 764 test hours. For attribution, the authors remove features in a set sequence, retrain, and assign each removed feature the resulting loss of out-of-sample R2. Bootstrap intervals and Benjamini-Hochberg false-discovery-rate (FDR) control determine which losses count. On that measure, timing (is_news_hour) leads; bank-specific activity counts help, LLM sentiment helps as a group, and aggregate news counts hurt.
The level forecast loses
At 1 hour, the best model has R2 of 0.9150 and NRMSE of 0.00549. The random walk reaches 0.9988 and 0.00065, roughly eight times less error. At 24 hours, the comparison is 0.8442 and 0.00742 against 0.9764 and 0.00289. The authors acknowledge that none of their machine-learning models beats the no-change random walk on level-based metrics.
They argue that the gap shrinks at longer horizons and that the benchmark loss "does not preclude attribution within the fitted models." Both points hold. Still, each FDR-tested attribution figure measures a change in R2 within a model already beaten by the naive forecast. Persistent price levels also make R2 forgiving. A gain of 0.012 leaves the model behind the no-change forecast, just by a little less.
What does the news-hour flag buy?
Removing is_news_hour costs a mean 0.012 of R2, with a 95% interval of [+0.006, +0.018] and q of 0.003. Four targeted activity measures pass FDR at +0.003 to +0.005. Two aggregate counts work against the forecast (news_articles_24h at -0.010), as does intraday_volatility at -0.010. WTI and the rate spread top conventional model-importance rankings yet contribute little when removed; the low-ranked flag produces the largest ablation loss. That reversal is a worthwhile finding.
Its practical reach is much smaller. Removing the flag changes directional accuracy by +0.04 points, and removing the whole timing group changes R2 by +0.0005. LLM sentiment alone changes it by +0.0020, news volume by -0.0018, and the combined group by +0.0001.
The sequence matters: each ablation credits a feature for what it adds after earlier removals. The authors acknowledge that dependence, and we did not find the removal order itself stated. Feature p-values pool stepwise ΔR2 across model-horizon combinations. The bootstrap also resamples window differences that are serially correlated, making the intervals and p-values entering FDR likely too narrow. These results do not test a rule that trades into news hours. The authors never ran one.
Direction, before trading costs
Directional accuracy reaches 66.08% at 1 hour, versus 63.05% for majority class and 62.61% for persistence. At 24 hours, it is 63.22% versus 51.14% and 57.70%. AUC falls from 0.744 to 0.634.
The class balance needs care. The paper forward-fills "missing observations, including those during weekend market closures," into a continuous grid; 17,517 hours covers almost every hour in two years. Only a strictly positive change counts as up, leaving each flat filled hour labelled down. This plausibly accounts for the 1-hour test set's 1,602 down moves and 939 up moves, while raising the majority baseline to 63.05%.
The authors report balanced accuracy of 0.721 at 1 hour and say the gains are "not attributable to class imbalance alone." The figure supports that claim. Given 1,602 down moves and 939 up, balanced accuracy above the 66.08% hit rate implies better classification of up moves than down moves. A classifier mainly collecting flat, forward-filled hours labelled down would behave the other way around. Table IV leaves one uncertainty: it selects the strongest result at each horizon, so hit rate and balanced accuracy may belong to different models. Before crediting the 3.0-point advantage over majority class at 1 hour, we would want accuracy recalculated for open-market hours.
The authors say their study "does not consider transaction costs, market frictions, or trading profitability." A tradable edge would have to fall in hours when USD/CAD moves far enough to cover the spread. They provide no size for the moves classified correctly.
Step 7 and the same rolling windows
The probabilistic comparison is striking. Ablation-selected Step 7 LightGBM reduces pinball loss by 55.6% at 1 hour (0.003329 to 0.001479) and 10.8% at 24 hours. Step 7, however, was selected through ablation on the same rolling windows used to score it; we found no separate holdout. The comparison also gives the full feature set its best model at each horizon while restricting Step 7 to LightGBM, a choice that disadvantages Step 7.
The uncalibrated intervals cover far too little. Full-feature 90% intervals achieve 53.72% to 63.23% coverage. Step 7 lifts 1-hour coverage from 53.72% to 68.92%, then falls to 46.29% at 24 hours, versus 57.85% for the full set. The authors distinguish sharpness from interval calibration. They describe conformal calibration, though we did not find calibrated coverage reported.
The underlying text remains sparse: 927 articles across 17,517 hours. A manual check of 100 articles is described as "generally consistent," with no agreement rate; the authors also offer an exploratory event study. Another discrepancy remains in the reporting. The text says baseline R2 exceeds 0.86 at every horizon, but the 24-hour table row says 0.8442. Nor does the paper explain how 2,541 one-hour test labels relate to 764 test observations per window. We found no Diebold-Mariano or Clark-West test against the random walk; the authors identify that work as a future step. We raised the same gap in a nine-series FX volatility comparison.
We could not run the study ourselves. We have no USD/CAD spot or FX futures data. A Canadian-dollar ETF trades during US equity hours on a different instrument, so substituting it would test another market.
Directional accuracy recalculated on open-market hours alone could change my verdict. For now, the flag has an FDR-supported mean R2 effect of 0.012 within a model that loses to a no-change forecast.