Model evaluation
There is one question any football model has to answer before it deserves attention: is it better than the forecast implied by the market? The closing line — the market consensus immediately before kick-off — is the reference standard in sports forecasting, and we measured ourselves against it. We publish no odds and no betting advice: the market appears here purely as a yardstick.
The verdict, over 504 matches across eight seasons: from matchday 14 onward the model matches the closing line, and edges slightly ahead of it. The whole deficit is concentrated in the first third of the season, when there are still few matches to learn from. Across the season as a whole the gap is +0.0010 with a standard error of 0.0021 — indistinguishable from zero.
We are writing this with the error bars in plain sight on purpose. The only difference on this page that survives its own standard error is the early-season one; everything else is consistent with a tie, and it would be dishonest to sell it any other way.
RPS difference between the model and the closing line, match by match. Negative = the model errs less. 504 matches, 8 seasons.
Matchdays 6 to 30
+0.0010
standard error 0.0021 · n = 504
The gap is smaller than half its own standard error: indistinguishable from zero.
Matchdays 6-10
+0.0098
standard error 0.0045 · n = 144
The market is ahead, and here the gap survives its standard error (t = 2.16).
Matchdays 14-30
−0.0025
standard error 0.0022 · n = 360
The model matches the closing line. The sign favours it, but stays inside the noise.
Forecast error (RPS) for the model and the closing line at each reference matchday. Lower is better. The two lines cross between matchday 10 and matchday 14.
Vertical axis truncated (0.165–0.205) so the crossover is visible; the real differences are the ones in the lower strip. Each point is one predicted matchday in each of the 8 seasons (n = 72 matches per point).
| Matchday | Model | Market | Difference | Std. error |
|---|---|---|---|---|
| 6 | 0.1874 | 0.1784 | +0.0090 | 0.0067 |
| 10 | 0.1888 | 0.1782 | +0.0106 | 0.0062 |
| 14 | 0.1838 | 0.1880 | −0.0041 | 0.0049 |
| 18 | 0.2016 | 0.2017 | −0.0002 | 0.0047 |
| 22 | 0.1719 | 0.1788 | −0.0069 | 0.0041 |
| 26 | 0.1779 | 0.1790 | −0.0011 | 0.0046 |
| 30 | 0.1843 | 0.1844 | −0.0001 | 0.0065 |
A note on uncertainty, because here it decides everything. The overall gap is +0.00103 with a standard error of 0.00207: the interval runs comfortably through zero, so we do not claim the market is ahead of us across the season. The only gap that survives its own standard error is the early-season one.
Model and market almost always see the same match. These are the cases where they did not.
In 53 of the 504 matches (10.5%) model and market named different favourites. In those, the market's pick won 22 times and the model's 15, with 16 landing on a third result. The error gap on that subset is +0.0101, with a standard error of 0.0079 — far too noisy to settle the question.
The eight matches where the two favourites were furthest apart. “Neither” means the third result came in.
Evaluated on 504 matches: eight seasons (2017-18 to 2024-25) × seven reference matchdays (6, 10, 14, 18, 22, 26, 30). At each point the model is fitted only on matches played up to that matchday and forecasts the next one — it never sees the future. Market probabilities are derived from closing odds published by football-data.co.uk (Pinnacle), with the margin removed using Shin's method, and cover 100% of the evaluated matches. The evaluated model is the same one that produces the forecasts published on this site.