A nineteen cents on the dollar longshot loss makes a strong headline. It is also the paper's weakest figure. The 95 percent interval stretches from -46.99% to +8.30%, and the Webb bootstrap null-imposed p-value is 0.184. Remove the ten largest parent events and the loss falls to -10.6%.

The authors acknowledge how much the answer depends on contract grouping. Their abstract says the findings "provide evidence of an aggregate FLB on Polymarket, but its magnitude depends on how contracts are grouped." FLB is the favorite-longshot bias. Magnitude is their word, though their own event-weighted panel changes direction and reaches +4.09%. What holds up is the scale of the tail mispricing: the pooled longshot probability error is -0.23 pp, and nowhere within the longshot tail does its absolute value exceed 0.53 pp.

The dataset Cardozo and Rivero-Wildemauwe assembled

The Polymarket Users dataset v1.3 (Grégoire, built on Akey and coauthors) runs from 11 November 2022 to 29 March 2026. It contains 588,287,492 observed transactions from 2,480,104 wallets, across 729,133 child markets nested within 316,429 parent events. Each child market is a separately traded question. A parent event is Polymarket's heading for related questions, such as the 2024 presidential winner with one contract for each candidate.

The filters come next. One excludes buyers whose counterparty concentration is at or above 0.5. It removes 17.7 percent of wallets and only 3.9 percent of dollars. The remaining sample has 586.1 million purchases worth $23.67 billion. By the cutoff, 560,909,624 purchases worth $22.49 billion had resolved. Every return reported below uses those 560.9 million resolved purchases.

For each purchase in a resolved two-outcome market, the authors calculate the terminal return on the assumption that the token is held through resolution. Purchases below 10 cents are longshots; those at or above 90 cents are favorites. They then use three aggregations: equal weight for every child market, pooling across all dollars, and pooling within each parent event before assigning equal weight to events.

The aggregation work is the part I would keep. The paper claims a second contribution as well, distinguishing persistent demand for the tails from who bears the resulting returns. The wallet evidence below supports that contribution.

Three weightings answer different questions

With equal weight per child market, longshots return -6.30% [-8.38, -4.22], while favorites return +0.277% [0.242, 0.312]. Pooling dollars produces a 19.35% longshot loss and a favorite return of +0.83% [0.60, 1.07]. Once every parent event receives one vote, longshots earn 4.09% [1.29, 6.90]. Favorites remain positive at +0.392%.

Moving from equal-child-market weighting to equal-parent-event weighting adds 10.392 percentage points. Reweighting across events accounts for almost the entire move, 10.329 pp. Pooling purchases inside each event adds just 0.063 pp. Weighting those events by longshot dollars then subtracts 23.440 pp.

Events containing many child markets therefore tend to have low longshot returns. The events receiving the greatest longshot spending also tend to have low longshot returns. These are separate concentration effects.

The right weighting follows the question being asked: what the crowd earned in dollar terms, whether a randomly selected listed question is mispriced, or whether the pattern appears broadly across events. The authors treat parent-event weighting as secondary because events combine mutually exclusive groupings with non-exclusive ones. Their caution is warranted. The 23.440 pp removed through dollar reweighting separates a pricing rule from concentration, while the event-weighted result remains positive.

Sorting events by child-market count gives equal-child-market longshot returns of -2.34% for single-market events, +39.51% for 2 to 5, -16.05% for 6 to 10, -15.40% for 11 to 25, and -34.16% for 26 or more. Event size has no monotone relationship with returns. Multi-market events still pull down the equal-child-market average.

Sports gives the game away

Crypto and Politics show the familiar two-sided pattern under both primary weightings. Crypto longshots lose 14.84% with equal weighting and 12.63% when pooled. Politics longshots lose 16.34% and 46.15%. All four corresponding favorite returns are positive.

Sports breaks the pattern. Its longshots earn +2.43% equal-weighted and +18.11% pooled, although neither Sports longshot interval excludes zero: [-0.93, 5.79] and [-74.68, 110.91]. More damaging is the equal-weighted favorite return, which is negative at -0.230%. Its [-0.285, -0.174] interval excludes zero. This is the load-bearing result. Sports has the most listed markets, averaging 5.00 child markets per event, with a median of 3 and a max of 144. A negative favorite return in that category strains any explanation centered on how people interpret probabilities.

Weather also reverses direction, from -25.15% equal-weighted to +24.74% pooled. Finance moves the other way: -2.62% equal-weighted versus -72.40% pooled. Only Crypto and Politics have all four intervals excluding zero. Culture and Tech run in the same direction, though their pooled longshot intervals contain zero. Anyone reusing these estimates takes that limitation with them.

Tiny errors in the tails

This is the result worth retaining. Probability errors, defined as payoff minus purchase price per token, are small at the extremes and much larger through the middle. Under equal weighting, they reach -5.86 pp near the 35-cent bin and +7.28 pp near the 60-cent bin. Within the longshot tail, the absolute error never exceeds 0.53 pp. The pooled longshot error is -0.23 pp, with a [-0.55, +0.09] interval that includes zero. For favorites, the error is +0.82 pp.

Division creates almost all of the headline -19.35%. Buyers paid $273.6 million for 23.4 billion longshot tokens, averaging 1.17 cents per token. Divide a quarter of a percentage point of error by 1.17 cents and the result is twenty percent. Prospect-theory explanations predict the greatest distortions at the extremes. Yet these prices are most accurate precisely there. As the authors put it, misperception fits the interior of the price range better than the tails where the bias is usually measured.

The trading case fails before costs enter the discussion. Capturing a 0.23 pp edge would require selling 1-cent tokens. The paper supplies no order-book depth or capacity estimates against which to size the position, and every reported return excludes fees by construction.

That is no trade.

No designated loser appears

Every month, wallets are ranked by the residual share of purchase dollars allocated to each tail. The residual comes from a regression of that share on volume, active months, market count, and Politics/Sports/Crypto shares during the prior six months. The top decile forms the recurrent group.

Demand is persistent. The top longshot decile comprises 10 percent of classified wallets and provides 26.6 percent of next-month longshot dollars. Among wallets still classified six months later, 30.4 percent remain in the same decile. Loss incidence does not track that persistence. The group earns -21.96%, compared with -20.33% for everyone else, and bears 22.6 percent of gross longshot losses despite supplying 26.6 percent of the dollars.

Below proportional.

For favorites, the recurrent decile contributes 15.1 percent of dollars and earns +0.25%, against +0.88% for other wallets. It captures 4.8 percent of net gains.

A risk-love account requires an identifiable clientele willing to pay for skew. The results identify none. Experience fares no better as an explanation. All five experience quintiles lose on longshots, from least to most experienced: -38.34%, -44.13%, -23.68%, -22.29%, -21.55%. The most experienced group outperforms its own market-week-price comparison groups by +0.83 pp [0.22, 1.44], while still losing in absolute terms.

Two more negative findings deserve attention. Longshot losses are largest in the heaviest-traded quartile, at -28.72%, compared with -4.08% in the thinnest. Shared collateral reduces the cash required to short mutually exclusive outcomes and is associated with larger losses: -25.76% pooled versus -9.79% without it. Both results are cross-sectional comparisons. The authors explicitly qualify the collateral result as "a cross-sectional difference rather than a causal estimate". For the trading-volume quartiles, the qualification appears in the appendix regression note.

The strangest figure concerns maker-side longshot purchases. They rise over short horizons, then fall by resolution. In the common sample, the marks are +4.85% at five minutes, +6.96% at an hour and +6.62% at a day, followed by -36.27% [-60.88, -11.66] at resolution. Those short-horizon figures use last transaction prices from five-minute bins without a trade-initiation sign. They do not represent prices at which anyone could necessarily have sold.

Why we left it untested

We hold spot crypto tokens rather than USDC-denominated binary event contracts, and we have no Polymarket trade records. We have previously used Polymarket prices as a benchmark in our note on esports rating models, where we read the quote instead of trading the book.

An aggregate bias based on 560.9 million resolved purchases changes sign when the aggregation unit changes. The paper's own event-level estimate is positive at +4.09%. A future dataset containing order-book depth could change my view if it showed tail probability errors of several points rather than tenths of a point. On the evidence here, the extreme-tail mispricing is too small to hold.