The Ledger
Every forecast this site publishes is frozen at publication and scored against the price the betting market closed at. This is the record. It is not a flattering one.
20,978 priced games scored · football 1888-89–2025-26 · NFL 1920–2025 · built 2026-09-02
The market wins, and this is by how much
Zero is a tie. Negative means the market was closer.
How this is measured
A model that beat a liquid closing market over twenty thousand games would be a business, not a website. What these two rows say is that our models get within a few percent of a price set by everyone in the world with money on the outcome, using nothing but the match results this site already holds — and that on the seasons below, they got closer still. The market is the benchmark precisely because it is hard to beat.
When the model said 70%, did it happen 70% of the time?9 bins
When the model said 70%, did it happen 70% of the time?
A well-calibrated forecast lands on the diagonal.
How this is measured
Green means the outcome happened more often than the model expected; pink means less. The bars are drawn to a six-point scale, so a full bar is a six-percentage-point miss.
The seasons the model beat the market9 seasons
The seasons the model beat the market
Derived, not chosen: every season the model's Brier came in under the market's.
How this is measured
This season, as it grades8 ledgers
This season, as it grades
Each hub freezes its call when it publishes it.
How this is measured
The market is four markets, and they disagree4 books
The market is four markets, and they disagree
Every posted price we can read on the same NFL game, and how each one leans.
How this is measured
Get the data: every book’s price on every game, raw. Each entry carries the prices as posted, the de-vigged number, and the book’s own hold, so the arithmetic here can be checked rather than taken.
48 of 48 games in the next 21 days carry two or more books. Leans are in percentage points at an even game, which is the one place a log-odds difference reads unambiguously in points.
Your own calls land on this axis too
Your own calls land on this axis too
Reader picks are scored on the same axis.
How this is measured
Call the matchweek before the model's card is revealed, rank your confidence, and take a side on the Upset Radar — the games where the model and the market disagree most. The model plays its own card, graded by the same rules, and so does the market.
Where these numbers come from, and what they cannot tell you
Why a skill score and not a Brier score
English football is a three-way question and a good model scores near 0.60 on it. The NFL and college football are two-way and a good model scores near 0.22. Those two numbers do not belong in one column, and putting them there would be exactly the kind of quiet wrongness this page exists to argue against. The skill score divides the model by the market that priced the same games, which removes the scale along with the question's shape.
What was priced, and by whom
Football odds are football-data.co.uk, available on 24 seasons of 127 — everything before that has no market to be measured against. Of those, 5,316 matches across 14 seasons carry a true closing price, which football-data has published only since 2012-13. The other 3,800 are scored against the pre-match price, because no closing price for those matches exists anywhere. The two are not interchangeable: a closing line has absorbed the team news and the money, so it is the sharper benchmark and the harder one to beat. This site scored the whole football sample against pre-match prices while calling them closing until 2026-08-30; correcting it moved the football skill score down, which is the direction an honest correction was always going to go.NFL odds come from covers.com, loaded as real numbers rather than inferred, after an earlier attempt to repair them by inference proved half right and therefore wrong. Live hubs read football-data for the Premier League and ESPN's DraftKings feed for the NFL and college football. No column anywhere on this site comes from an exchange yet; a Kalshi or Polymarket price would be a genuine fourth column, not a replacement for any of these.
What this does not prove
Nothing here is held out. Both historical models were fitted on their full histories, so the calibration above is in-sample and should be read as a consistency check rather than as evidence of predictive skill. The football model has almost no skill before about 1960 — 3.2% over an era baseline across the whole series, and a fraction of that in the early decades. Twenty fixtures in the football source are recorded the wrong way round and two scorelines are off by one; they are listed rather than repaired, because sourcing the real result is the only honest fix.
Why the elections row is empty
A seat range cannot be scored until the seats are counted. The election forecast publishes its own history of what it said over time, but it has no resolved outcomes to be graded against until November, and inventing an accuracy figure for it would undermine every other number on this page.
Sources
Against Expectation carries the full historical ledgers and their reconciliations. The live hubs are Premier League, NFL, College Football and MLB. Methodology for each model is stated on its own hub.