The Ledger

Every forecast this site publishes is frozen at publication and scored against the price the betting market closed at. This is the record. It is not a flattering one.

20,978 priced games scored · football 1888-89–2025-26 · NFL 1920–2025 · built 2026-09-02

Priced games
20,978
scored against a market price
Seasons won
9 of 71
model closer than the market
Calibration sample
51,330
9 probability bins
Graded this season
105
and counting

The market wins, and this is by how much

Zero is a tie. Negative means the market was closer.

How this is measured
Brier scores are not comparable across a sport with draws and one without, so the unit here is the skill score against the market that priced the same games: one minus the model's Brier over the market's.

A model that beat a liquid closing market over twenty thousand games would be a business, not a website. What these two rows say is that our models get within a few percent of a price set by everyone in the world with money on the outcome, using nothing but the match results this site already holds — and that on the seasons below, they got closer still. The market is the benchmark precisely because it is hard to beat.

When the model said 70%, did it happen 70% of the time?9 bins

When the model said 70%, did it happen 70% of the time?

A well-calibrated forecast lands on the diagonal.

How this is measured
Every English top-flight match outcome the model has ever priced, sorted into probability bins: the share that actually happened should match the share it was given. 51,330 outcomes across 9 populated bins.
0.0-0.1
predicted 8.6% · actual 3.6%
5.0
28 n
0.1-0.2
predicted 16.4% · actual 15.4%
1.0
765 n
0.2-0.3
predicted 25.8% · actual 25.2%
0.6
2,378 n
0.3-0.4
predicted 35.9% · actual 37.7%
+1.8
6,557 n
0.4-0.5
predicted 45.4% · actual 46.2%
+0.8
14,789 n
0.5-0.6
predicted 54.8% · actual 55.5%
+0.7
16,471 n
0.6-0.7
predicted 64.1% · actual 63.1%
1.0
8,207 n
0.7-0.8
predicted 73.4% · actual 76.7%
+3.3
2,036 n
0.8-0.9
predicted 81.6% · actual 84.9%
+3.3
99 n

Green means the outcome happened more often than the model expected; pink means less. The bars are drawn to a six-point scale, so a full bar is a six-percentage-point miss.

The seasons the model beat the market9 seasons

The seasons the model beat the market

Derived, not chosen: every season the model's Brier came in under the market's.

How this is measured
These are the exceptions, and naming them is the only way the rest of this page means anything. A page that only ever showed its own scoreboard would be worth nothing.
English top flight · 1 of 24
  • 2008-090.5678 v 0.5725
NFL · 8 of 47
This season, as it grades8 ledgers

This season, as it grades

Each hub freezes its call when it publishes it.

How this is measured
Nothing below is back-filled, and a row with nothing in it says so rather than waiting for a good week to appear.
The market is four markets, and they disagree4 books

The market is four markets, and they disagree

Every posted price we can read on the same NFL game, and how each one leans.

How this is measured
Until now 'the market' on this site meant whatever single price ESPN was carrying, which is DraftKings. That is a market the way one poll is an electorate. This board prices each game at four books, removes each book's margin by the power method rather than proportionally (proportional de-vig leaves longshots too high, because books load their margin onto longshots), and averages what is left in log-odds. A book's lean is measured against the consensus of the OTHER books on the same game, never against a consensus that includes itself: including it drags every book's own baseline toward its own number and shrinks the very thing being measured. Positive means the book sits higher on the home side than its peers. A price we had to translate from a spread is shown but never votes, and never earns a lean.
DraftKingssportsbook
48 priced · no lean yet
Polymarketexchange
48 priced · lean -0.51pp
FanDuelsportsbook
16 priced · lean +0.96pp
Kalshiexchange
16 priced · lean -0.46pp

Get the data: every book’s price on every game, raw. Each entry carries the prices as posted, the de-vigged number, and the book’s own hold, so the arithmetic here can be checked rather than taken.

48 of 48 games in the next 21 days carry two or more books. Leans are in percentage points at an even game, which is the one place a log-odds difference reads unambiguously in points.

Your own calls land on this axis too

Your own calls land on this axis too

Reader picks are scored on the same axis.

How this is measured
Citizen of Nowhere Picks scores a reader's hard calls with the same Brier the 1920 ledger uses, so a pick made this Saturday is measured the way a match from 1958 is measured.

Call the matchweek before the model's card is revealed, rank your confidence, and take a side on the Upset Radar — the games where the model and the market disagree most. The model plays its own card, graded by the same rules, and so does the market.

Where these numbers come from, and what they cannot tell you

Why a skill score and not a Brier score

English football is a three-way question and a good model scores near 0.60 on it. The NFL and college football are two-way and a good model scores near 0.22. Those two numbers do not belong in one column, and putting them there would be exactly the kind of quiet wrongness this page exists to argue against. The skill score divides the model by the market that priced the same games, which removes the scale along with the question's shape.

What was priced, and by whom

Football odds are football-data.co.uk, available on 24 seasons of 127 — everything before that has no market to be measured against. Of those, 5,316 matches across 14 seasons carry a true closing price, which football-data has published only since 2012-13. The other 3,800 are scored against the pre-match price, because no closing price for those matches exists anywhere. The two are not interchangeable: a closing line has absorbed the team news and the money, so it is the sharper benchmark and the harder one to beat. This site scored the whole football sample against pre-match prices while calling them closing until 2026-08-30; correcting it moved the football skill score down, which is the direction an honest correction was always going to go.NFL odds come from covers.com, loaded as real numbers rather than inferred, after an earlier attempt to repair them by inference proved half right and therefore wrong. Live hubs read football-data for the Premier League and ESPN's DraftKings feed for the NFL and college football. No column anywhere on this site comes from an exchange yet; a Kalshi or Polymarket price would be a genuine fourth column, not a replacement for any of these.

What this does not prove

Nothing here is held out. Both historical models were fitted on their full histories, so the calibration above is in-sample and should be read as a consistency check rather than as evidence of predictive skill. The football model has almost no skill before about 1960 — 3.2% over an era baseline across the whole series, and a fraction of that in the early decades. Twenty fixtures in the football source are recorded the wrong way round and two scorelines are off by one; they are listed rather than repaired, because sourcing the real result is the only honest fix.

Why the elections row is empty

A seat range cannot be scored until the seats are counted. The election forecast publishes its own history of what it said over time, but it has no resolved outcomes to be graded against until November, and inventing an accuracy figure for it would undermine every other number on this page.

Sources

Against Expectation carries the full historical ledgers and their reconciliations. The live hubs are Premier League, NFL, College Football and MLB. Methodology for each model is stated on its own hub.