Stable team identity carries nearly all the signal recovered from Guan and Tung's 5,135-game corpus. On the 912-game held-out split, removing dynamic form from their own two-stage model raises Brier from 0.2257 to 0.2268, a cost of 0.0011. Remove stable strength and the model falls back to the purely dynamic logistic benchmark at 0.2351, another 0.0083 behind.

The "dynamic" architecture gets little from its dynamic component.

Guan and Tung make the same point. The two blocks contribute very different amounts, with stable team identity carrying more recoverable signal than recent form. They retain the form block anyway.

The model they built

The target is a pre-game win probability for a single map, evaluated by ranked probability score. With a two-outcome contract, ranked probability score collapses exactly to the Brier score. A constant 0.5 forecast produces 0.25.

The proposed model stays deliberately small. It has Four fixed terms: an intercept, a first-pick indicator, and two side-specific slopes. The slopes apply to an exponentially weighted moving average of each team's past results on that side, which supplies the form term. First pick comes from the schedule rather than being earned. A per-team strength block sits above those terms, with one coefficient per team pulled toward zero through a ridge penalty. Four coefficients and the shrunk team block are fitted in one pass on win/loss log-loss.

The paper proves that the ridge penalty gives exactly the maximum-a-posteriori solution for a logistic mixed model whose team strengths follow a normal prior. Its weight represents the inverse prior variance. The chosen value implies a strength standard deviation of 0.5 on the log-odds scale.

The comparison model is the classical two-stage route, also built by the same authors. Its first stage fits a Gaussian mixed model to a fixed composite response: 0.50 outcome plus weighted, standardised gold, kill and tower margins. Restricted maximum likelihood supplies the variance components. Team strengths are shrunk predictions under the fitted prior, and a Platt map converts the final ratings into probabilities.

The corpus contains 5,135 games with populated first-pick assignments from 5,893 collected. LPL (1,822) and LCK (1,276) provide most of them. Four additional regional leagues and three international events supply the remainder: Worlds 190, MSI 157, First Stand 68. The sample covers roughly 80 teams from 2024 to 2026. Two evaluation protocols appear throughout. One uses a global 80/20 time split. The other refits at every game in walk-forward fashion, subject to a 500 games minimum.

The comparison survives

The paper re-implements the classical dynamic Bradley-Terry model of Cattelan, Varin and Firth on the same data. Both paper models beat it comfortably. The strongest Cattelan configuration records 0.2351 on the holdout, trailing by 0.0121 and 0.0094. For the two-stage gap, the paired standard error is 0.0052 (t = -1.82, p = 0.07). The 4,605-game walk-forward run preserves the ordering, with gaps of 0.0112 and 0.0104.

The authors openly tune the baseline's decay by oracle on the evaluation set itself. Its 0.2351 therefore receives favorable treatment. A static Stefani fit without form terms scores 0.2301, while the shrunk version reaches 0.2268. In this corpus, a time-invariant rating forecasts better than a running form score.

Most remaining choices barely move the result. On a later 4,703-game corpus, using the selected penalty, walk-forward Brier shifts only 0.0002 across the full decay grid for the form average, 0.1 to 0.9. Every dense re-weighting of the two-stage composite response returns 0.2257 to four decimals, including the version that drops the win/loss term outright. Correlations among the three margins run from 0.89 to 0.96, leaving the weights with little influence.

Adding richer end-of-game features to the latent state steadily reduces accuracy: 0.2230 at d = 1, 0.2242 at d = 4, and 0.2258 at d = 10. The end-to-end difference is 0.0028, with standard error 0.0020 and p = 0.16. The ordering supports a mechanism claim, though nobody should describe it as significant. Its source appears in the weights. As correlated margins enter, the learned coefficient on the win indicator flips to -0.81 and then -0.88.

Estimation makes the case

Accuracy leaves the one-stage model without a clear win, as the paper acknowledges. On the holdout, it scores 0.2230 against 0.2257, a paired difference of +0.0027 with p = 0.22. Walk-forward results on 4,605 games are 0.2207 against 0.2215, giving +0.0008 and p = 0.40. The market-window difference is +0.0003. For the 10-feature one-stage variant it is -0.0001 (p = 0.97). Per-game scores from the two models correlate at 0.944.

Their rating blocks also rank teams almost identically. Spearman correlation with realised test win rates is 0.438 for the ridge strengths and 0.436 for the shrunk mixed-model predictions. The calculation covers 50 teams with at least five test games.

The argument rests on estimation hygiene. The one-stage fit is convex and converges on thin slices where the classical variance-component surface becomes singular. Within-league LCP provides the example: 0.2648 against 0.3055 across 86 test games. On the 912-game holdout, its calibration slope is 0.88. The two-stage model reaches 0.67 on the same holdout even though its Platt map was fitted in sample.

Across the full walk-forward run, the proposed model is essentially perfectly calibrated, with intercept -0.009 and slope 0.995, despite having no calibration step. The authors flag inflation from the in-sample Platt fit and partially repair it. Refitting out of fold moves 0.2257 to 0.2248.

Where Polymarket wins

Across 924 maps matched from October 2025 onward, the market leads by 0.009 Brier (t = +2.24, p = 0.025). On 136 maps, however, the price comes from the series-winner contract at the decider point, when the map and series represent the same bet. Removing those maps reduces the gap to +0.006 (t = +1.41, 95% CI -0.002 to +0.015). The authors interpret this as bounded parity on per-game contracts. The bounds deserve emphasis: any market edge below 0.015 and any model edge below 0.004.

The proposed model records a walk-forward Brier of 0.2263 during that window. The paper's other route scores 0.2260 on the same maps. Against the constant forecast at 0.25, the model gains 0.024. With deciders included, the market's 0.009 advantage is about 40 percent of that gain. Removing them leaves 0.006, or 25 percent.

Worlds accounts for the sharpest deficit: +0.069, t = 4.61, across 50 maps. It is the only league slice that survives a Bonferroni correction across eight. Deciders supply the other concentrated weakness, and their prices represent logically identical claims rather than imputations. The ridge block has no patch term and no player-level information across roughly 80 patches. An omission of that kind could produce a cross-region deficit this large.

The filter creates another weakness by requiring a team to have appeared previously on the relevant side. On the cross-region holdout, 115 of 1,027 test rows disappear because a team has yet to play on its assigned side. Without filtering, the two-stage model worsens to 0.2312. Its score is near 0.27 on those fallback games and near 0.23 elsewhere.

Scoring rules are the limit

The paper confines its Polymarket comparison to forecasting quality, model RPS against market-implied RPS, and discusses nothing built from those probabilities. That boundary should be taken literally. Prices are the final pre-start quotes, with median staleness of 0.5 minutes. A quote does not establish an executable fill. The analysis says nothing about depth at that quote, fees, position limits or settlement timing. A 0.006 Brier gap against one final pre-start price offers thin cover for any of them.

We could not test any of this ourselves. Our data lack the map-level schedule, side assignment, first-pick rights and outcomes. They also lack historical Polymarket esports order books.

A decider-and-Worlds subsample three or four times the size of these 50 and 136 maps would change my mind on the market comparison, provided pricing came from the book rather than one quote. The lasting contribution is measurement. On this 912-game holdout, stable ratings carry the signal, while the form block contributes 0.0011 Brier.