EdgeBoard
Open the board

← All articles

Calibration, Brier score and log loss, in plain words

Accuracy is the first thing people ask about a forecaster and the least informative. Three better measures, explained with arithmetic you can do on a napkin.

September 22, 2026 · 3 min read · Agent Creative

Ask how good a forecaster is and you usually get one figure: the share of games it got “right”. It is the easiest number to understand and the easiest to mislead with. Here are the measures that tell you more, and the two baselines every one of them should be set against.

Why accuracy alone is not enough

Accuracy treats a forecast as a yes or a no. If the model gave the home team 51% and the home team won, that is a hit. If it gave them 95% and they won, that is also a hit. The two forecasts could hardly be more different, and accuracy cannot tell them apart.

It also has a ceiling nobody mentions. If a league’s games are mostly close, even a perfect forecaster, one whose probabilities are exactly right, will be “wrong” about four times in ten. Low accuracy can mean a bad model or an unpredictable sport. You need other measures to know which.

Calibration: do the numbers mean what they say?

Sort the games into groups by what the model said: the 55% games, the 65% games, the 75% games. Then count how often the favorite won in each group. A calibrated model’s 65% group wins about 65% of the time.

A calibration chart plots exactly that: what was said along the bottom, what happened up the side. Points on the diagonal are good. Points below it mean overconfidence, because favorites won less often than promised.

One caution. A group with 40 games in it will wander from the line by luck alone. Even a perfect forecaster’s 70% group could easily come in at 62% or 78% on 40 games. A fair chart shows that range, and a point is only a problem when it sits outside it.

Brier score: the size of the miss

The Brier score measures how far each forecast was from what happened. Write a win as 1 and a loss as 0, subtract the forecast, square it, and average over all the games.

  • Say 70% and the team wins: (0.7 − 1)² = 0.09.
  • Say 70% and the team loses: (0.7 − 0)² = 0.49.
  • Say 50% on anything: 0.25, every time.

Lower is better. Zero is perfection, and 0.25 is what a coin flip scores. A sports model that gets to 0.23 or 0.24 over a season is doing honest work, because the games really are that uncertain.

Log loss: the price of being sure and wrong

Log loss is similar, but it punishes confident mistakes much harder. The score for a game is the negative logarithm of the probability you gave to what actually happened.

  • Gave the winner 70%: 0.357.
  • Gave the winner 50%: 0.693.
  • Gave the winner 10%: 2.303.

Again, lower is better, and a coin flip scores 0.693 on every game. A forecast of 99% that loses scores 4.6, about as much as five misses at 60%. Log loss rewards a forecaster for knowing how much it does not know, which is why EdgeBoard’s settings were tuned to minimize it.

The two baselines

None of these numbers means much alone. Set them against two forecasters that need no skill:

  • A coin flip. 50% on every game. Accuracy 50%, Brier 0.25, log loss 0.693.
  • Always pick the home team. Home teams win more than half their games in every major league, so this one is harder to beat than it sounds.

A model earns attention only when it beats both, on the same games, by more than luck would explain. When EdgeBoard’s settings were tuned on September 30, 2026, the held-out hockey season scored a log loss of 0.6892, against the coin’s 0.6931. That is a thin margin. It is on the Method page because leaving it out would make the model look better than it is.

How many games is enough?

More than feels reasonable. Over 20 games, a coin can call 14 by chance. Over a few hundred, the scores start to settle. Over a few thousand, calibration can be read group by group.

That is why a short record should be read as a short record. The Record page shows how many published predictions have been decided so far and says plainly when that is too few to judge. The sample size is part of the result.