Polymarket traders can trust the feed's taker side. On all 24.5 million matched pairs, the public print identifies the taker's side and token exactly. That gives Dubach a firm label for measuring what classifiers do to cost estimates. The effective-spread distortion survives on every sampled day. The abstract's volume-weighted impact reversal is less settled: remove tied prints, and the eligible sample falls from 16.4 million fills to 9.1 million.
The label comes from settlement
Polymarket matches orders off-chain, then settles on Polygon. Each public WebSocket trade print carries its settlement transaction hash. The transaction emits a fill leg for each order; exactly one leg names the exchange contract as counterparty, identifying the taker order. Dubach joins a third-party (PMXT) capture of the feed to that leg by hash across twelve selected days between 26 April and 5 August 2026. Two days precede the 28 April move from first-generation exchange contracts (V1) to the second generation (V2). V1 fill events lack a side field, so he reads the side from whether the maker supplied collateral or tokens.
The taker-leg distinction matters. A YES buyer can settle against a NO buyer through a mint, creating a complete set from collateral. Mints account for 53.5% to 73.8% of settled transactions. Counting every leg also counts those makers as buyers. The buy share is 80.6% to 85.6% for taker legs and 84.7% to 88.4% for all legs. In five-second cells, agreement with the feed falls to 86.8% to 90.9%, below the majority-class baseline of 0.906 to 0.947. Print matching covers 99.8% to 100.0% from May onward and 92.1% to 96.8% in April. Only about 72% of settlements on 5 June and 1 July have a print; the author considers collector faults likely.
Dubach then removes the side field and signs each print using Lee-Ready, the quote rule, the tick test and bulk volume classification (BVC). BVC gives a five-second bar a buy share based on its price change. Its run is retrospective because it needs bar closes and a whole-day fallback volatility. Lee-Ready and the quote rule take the last quote received strictly before the print. For each sign, he measures three cent-per-share outcomes on the same 16,448,261 fills: effective spread 2a(p-m), realised spread against the cached midpoint five minutes later, and the five-minute midpoint move between them. The distortion is exactly minus twice the weighted sum of taker-signed outcomes on classification errors. An error count alone cannot tell you the cost error; the location of those errors determines it.
Why does 0.942 still move the spread?
All four rules can sign 16,615,817 prints. Their equal-day balanced accuracies are 0.942 for Lee-Ready, 0.941 for the quote rule, 0.768 for tick and 0.670 for BVC. MCCs follow at 0.809, 0.799, 0.409 and 0.260. Lee-Ready wins on all twelve dates. Yet on the 16.4 million fills in the cost comparison, it adds 0.429 cents to equal-fill effective spread against the taker baseline of 2.328. The increase appears on every date.
Random sign errors would pull a signed spread toward zero. Lee-Ready moves it upward because it assigns side relative to the delivered midpoint, flipping takers who traded on its favourable side. On 17 July, a confirmed buy of 115.07 shares at 0.22 faces a delivered 0.22/0.24 quote. Lee-Ready marks it as a sell, changing that fill's effective spread from -2 cents to +2.
Realised spread can conceal the same problem. Lee-Ready's equal-fill mean difference is only -0.209 cents, although its average absolute per-fill difference reaches 6.786 cents. The individual errors offset one another in the mean. Share-volume weighting puts the difference at +0.365 cents.
Tied prints and the impact reversal
The taker benchmark changes with weighting: impact is -4.158 cents on equal-fill weights and +0.689 on share-volume weights. At share-volume weights, tick gives -0.374 and BVC -0.119. Their daily means disagree in sign with the taker on 8 and 6 of 12 dates.
Many prints carry the same asset and receipt millisecond. Their order comes from cache rows, while venue order is unknown. These ties make up 36.9% of matched prints. Exclude them from scoring and 9.1 million fills remain. Taker impact becomes 0.955, tick 0.102 and BVC 0.181, all positive; daily reversals drop to 5 and 2.
Dubach acknowledges the change in the abstract. The untied check, he says, neither restores venue order nor decides which population to prefer. I agree. The reversal remains undetermined, and I would not repeat it as a finding. Tick still distorts the estimate after the cut, at 0.102 against the taker's 0.955. In the full sample, its -1.063 distortion exceeds the +0.689 taker level and flips the sign. Without ties, the distortion is a gap of 0.853 against 0.955, short of a flip. A one-second quote-freshness screen leaves tick and BVC negative, with 9 and 7 flipped days, though Lee-Ready reverses one daily mean there.
Classification rank does not settle cost rank. The quote rule's equal-fill effective-spread error is 0.334 cents, below Lee-Ready's 0.429. BVC has the smallest volume-weighted realised-spread error of the four, at -0.016.
Quotes as delivered
The spreads use quotes when the collector received them. On a separate sample that every lag variant can sign, Lee-Ready's balanced accuracy falls from 0.9538 with the preceding quote to 0.885 at a one-second lag. The archive cannot control those clocks. Nor can it establish the age or two-sidedness of the five-minute midpoint: the saved midpoint has neither a timestamp nor a bid and ask. Fees and rebates are excluded. The author explicitly limits these figures: they are not costs against the execution-time venue book, profitability or causal impact.
We could not test these results ourselves. Our data contain no Polymarket outcome tokens, trade prints or quotes. One-minute bars cannot recreate a quote-strictly-before-print midpoint or assign a sign to each fill.
For Polymarket trading, use the feed's side field; it is exact in the matched data. The classifier exercise applies to archives with prints and quotes that lack that field. Without prints, even recovering the sign is harder. In the author's companion study, resting-size decrements recovered the on-chain sign in only about 59% of matched buckets.
Other venues should be wary of treating a 0.94 classification score as a licence to trust signed costs. Qin and Yang report 0.4983 tick raw accuracy on V1, versus 0.704 and 0.710 here, a difference of about 0.21. Their filtered Standard Binary sample spans the V1 lifecycle; these figures are settlement-filtered own-rule scores from two V1 days. Period, unit, sign convention and sample selection differ. The paper flags that it cannot reconcile the gap.
Venue-clock ordering of the tied prints would have to reproduce the impact reversal before I would call it a finding.