Brazilian rates traders should care about the gap between diagnosis and guidance. Across the 21 statements from 2024 through August 2026, mean tone is +0.237 and mean guidance is -0.024. The 24 statements from 2021 to 2023 have a nearly unchanged tone reading of +0.263, while guidance averages +0.458. The diagnosis therefore remains hawkish as the willingness to indicate the next move drops from +0.458 to a near-zero -0.024. In the paper's words, the recent window "remains strongly positive at +0.237, even though its average guidance score is near zero." A genuine separation would deserve space on a trading screen.
We did not test it, and could not have. An end-to-end reproduction requires the raw Copom statements, the LLM model identifier, the prompt version and the inference settings. The paper records that these are absent from the supplied artifacts. Our data also exclude Selic decisions and DI futures. Nothing below tests the method, and every reported figure belongs to the author.
The author disclaims a predictive result from the outset. The abstract says: "These are descriptive outputs, not a validated forecast of Selic decisions or DI returns." It then defines the intended contribution: "The main contribution is therefore methodological: a transparent, incremental, auditable system that separates rhetorical tone from policy guidance and uncertainty." The conclusion returns to the same boundary: "The project's most credible current contribution is measurement and organization, not prediction." There is little value in judging a forecast the paper never claimed. A narrower question remains. What can a rates desk learn from a measure whose only computed correlation is with another product of the same system?
Everything rests on the construction of the tone and guidance series.
How sentences become scores
de Macedo Santos turns the official Portuguese Copom statements into a table of dates and text. Any row without a date is removed. If a date occurs twice, the final observation is retained. A regular expression separates sentences at whitespace following a period, question mark, exclamation point, semicolon or colon, provided the next token begins with an uppercase letter, a number, a quote or a parenthesis. Segments shorter than four whitespace-delimited tokens are discarded. A single API call can contain up to forty segments. The Portuguese prompt assigns each segment to one of four mutually exclusive classes: hawk, dove, neutral, out of context.
The same prompt requests a sentence class and the short original expressions supporting a hawkish or dovish interpretation. Each expression receives a 0-to-1 intensity weight. The bands are 0.1 to 0.3 for weak, 0.4 to 0.6 for moderate and 0.7 to 1.0 for strong. Parsing is followed by clipping. Under one prompt instruction, a sentence that merely records the mechanical rate decision is neutral and contains no signals. This is meant to score the language surrounding the action without encoding the action twice. Three dates serve as calibration anchors: March 2020 as strongly dovish, May 2021 as hawkish, August 2023 as dovish.
At the document level, the mean intensity of all hawkish signals is multiplied by the number of hawkish sentences. The equivalent dovish product is subtracted. The difference is divided by the combined count of hawk, dove and neutral sentences, leaving out those classed as out of context, and clipped to [-1, 1]. If one side has no extracted signals, its intensity defaults to 1.
A separate call examines the full document and produces four discrete fields. Guidance direction takes (-1, 0, +1). Guidance explicitness takes (0, 0.5, 1), depending on how firmly the statement commits. Uncertainty level ranges from (0 to 3), while uncertainty change takes (-1, 0, +1). Multiplying direction by explicitness gives guidance, which remains outside the tone index by design. This product produced the -0.024 reported above.
The sample ends in August 2026
The sample contains 80 statements from August 31 2016 through August 5 2026 and 1,498 classified segments. Denominators include 1,400 of them; 98 are removed as out of context. Neutral leads with 631 sentences (42.1%), followed by hawkish at 499 (33.3%) and dovish at 270 (18.0%). The mean document score is +0.107, with a median of +0.106, a standard deviation of 0.208 and an interquartile range of -0.041 to +0.265. There are 26 negative readings and 54 nonnegative ones.
August 4 2021 is the most hawkish statement at +0.570, based on 14 hawkish and zero dovish sentences. January 11 2017 is the most dovish at -0.357, with two hawkish and nine dovish.
A rates desk would immediately compare the score with the Selic decision and the DI curve. The paper contains neither test and presents the output as a monitoring index rather than a trading signal. Itaú's iSent, the acknowledged inspiration and a bank sentiment index for Copom communication, reports correlations of 0.79 with the contemporaneous Selic change and 0.77 with the one-meeting-ahead change. The paper explicitly assigns those figures to iSent, and they must not be attributed to this index.
August 2026 across four fields
The tone reading is +0.232, drawn from eight hawkish, two dovish and nine neutral sentences. Guidance direction is 0 and explicitness is 0.5. Uncertainty level reaches 3, with change at +1.
The hawkish material concerns the risk diagnosis: upside-skewed inflation risks, unanchored expectations, a resilient labour market and exchange-rate depreciation risk. None commits the committee to a direction. Hence the zero direction and partial explicitness. The design proves useful here, since one scalar would have shown hawkishness while dropping the conditionality.
In my judgement, rather than the author's, the structural layer contributes less than four fields imply. Every one of the 80 statements has an uncertainty level of 2 or 3. Regime averages inch from 2.23 (August 2016 to 2020), to 2.54 (2021 to 2023), then 2.76 (2024 to August 2026). The categories themselves barely vary; only their group means shift, and they do so slowly. Uncertainty change equals zero in 64 of 80 cases, alongside 13 increases and three decreases. The author partly links this concentration to the instruction that the prompt return zero whenever the prior-meeting comparison lacks support.
Every statement receives some form of guidance. Direction carries much of the variation, with 32 dovish, 22 ambiguous and 26 hawkish readings, alongside the 44/36 division between full and conditional explicitness.
Counts scaled by a document-wide mean
For each class, the sentence count is multiplied by the document's mean intensity across all same-sign phrase signals. Sentence-level weights are never added individually. The author states the drawback directly. The rule "gives identical count weight to a strongly hawkish sentence and a weakly hawkish sentence while scaling the entire hawkish count by the document's mean signal intensity."
This choice has two effects. The intensity averages include every extracted signal, irrespective of the final class assigned to its host sentence. A dovish expression found inside a hawkish or neutral segment can therefore alter the dovish intensity term. The default intensity of 1 creates pressure in the other direction. When a statement contains dovish sentences without any parsed dovish expression, the dovish weight becomes 1.0, compared with the sample's mean dovish weight of 0.586.
Segmentation brings another limitation, one the paper acknowledges. Because scraped text is split by regex, missing spaces after punctuation can fuse distinct sentences. A merged segment containing both a hawkish and a dovish phrase still gets one class. The author flags this issue and recommends replacing the regex with a dedicated Portuguese tokenizer.
Document length controls the influence of any single label. January 2017 has 13 denominator sentences. Moving one of its nine dovish sentences into neutral changes the score by roughly 0.045, about a fifth of the full series' 0.208 standard deviation. Intensities are recalculated within each document for the live score. According to the paper, the global 0.654 and 0.586 averages serve as diagnostics and never enter the score. The 0.586 used here therefore substitutes for the January 2017 document mean. The arithmetic is mine, using the paper's counts. Appendix B reaches the same conclusion about the greater discreteness of earlier, shorter statements.
The implementation audit admits another leak. Although the prompt directs the model to treat decision-only sentences as neutral, some historical outputs break that rule. The proposed remedy is post-processing or a gold-set check. For an unknown portion of the sample, the mechanical rate move consequently enters the tone score.
What does 0.719 establish?
The paper computes only one correlation for its own index. Tone versus guidance produces a Pearson 0.719 and a Spearman 0.718. Both originate from LLM prompts run over the same documents. The paper expressly describes this relationship as descriptive and warns against treating it as independent validation. The remaining correlations discussed in the text, 0.79 contemporaneous and 0.77 one meeting ahead, belong to iSent and are identified accordingly.
About half of the variance is shared: 0.719 squared is 0.52. A quick reader may see confirmation that the construct captures something real. I see a consistency check on one annotator answering the same broad question twice.
The author's own accounting lists the missing work required before the index becomes tradable. The proposed validation table calls for an economist-labelled gold set, with precision, recall and a confusion matrix. It further requests repeated-run stability, an alternative model and prompt, an unweighted iSent-style benchmark, a Selic lead/lag study, and a DI event study that reports costs and drawdowns. A note beneath the table says the supplied artifacts contain none of these results. Accuracy has no numerical value without the labelled set because the same model provides the class and the weight.
The three calibration dates also fall within the study window. The author says they guide the model while leaving the mechanical score unconstrained, and reports that realized scores differ materially from the suggested anchors.
Contemporaneous correlation is the easier result to obtain. We made the same argument in our review of FinSMART (note). Evidence that would change my view begins with the one-meeting-ahead Selic correlation from a sample holding out the anchor dates, followed by a DI event study net of costs. Until those exist, the index remains what Appendix B calls one input to a broader communication dashboard. Its demonstrated separation amounts to the +0.237 tone against -0.024 mean guidance from 2024 to August 2026, alongside a tone-guidance Pearson of 0.719. Appendix B supplies the warning a trader should keep beside the reading: a score of +0.23 does not mean a 23% probability of a rate increase.