EdgeBoard
Open the board

← All articles

A backtest is not a track record

Both are tables of predictions and results. One was written down before the games and one after. Why that difference matters, and four questions to ask of any “proven” model.

September 15, 2026 · 3 min read · Agent Creative

Two tables can look identical. Each lists games, a probability for each, and what happened. One shows a model that was right 60% of the time last season. So does the other. The difference is when the probabilities were written down, and that difference is most of what matters.

What a backtest is

A backtest takes a model that exists today and runs it over games that have already been played. Done carefully, it is walk-forward: the model predicts each game using only the games before it, then learns the result and moves to the next. No game is predicted with knowledge of its own score.

That is a real and useful exercise. It shows how a method behaves across thousands of games, which a new model cannot show any other way. EdgeBoard runs one and publishes it.

How a careful backtest still flatters

Even a walk-forward backtest knows things about the past that nobody knew at the time.

  • The settings were chosen with hindsight. How fast ratings move and how big the home edge is are numbers picked because they worked best on past seasons. Applied to those same seasons, they look better than they will on new ones.
  • The model itself was chosen with hindsight. For every backtest you are shown, there were versions that did worse and were dropped. The survivor’s record includes the luck that made it the survivor.
  • The data is cleaner than it was. Postponed games, corrected scores and moved start times are all tidy in the archive. A live model has to cope with them as they happen.
  • Nobody was watching. A number that was never public was never exposed to being checked.

None of this makes a backtest dishonest. It makes it an upper estimate. The honest way to present it is with a label.

Holding a season back

One partial fix is to tune on some seasons and test on another the tuning never touched. EdgeBoard’s settings were tuned on three seasons and then checked on a fourth. For the NFL the two came out close: a log loss of 0.6327 on the tuned seasons and 0.6351 on the season held out, as recorded on September 30, 2026. For the NHL the held-out season was clearly worse, 0.6892 against 0.6617.

That gap is what hindsight is worth. It is on the Method page for each league, and it is the reason nobody should quote the tuned-season figures as if they were the model’s record.

What a track record is

A track record is the list of predictions that were made public before the games started and cannot be changed. It has three properties a backtest lacks.

  • A time stamp. Each prediction is stored before the start of its game.
  • No edits. Rows are added and never altered. On EdgeBoard each row carries a fingerprint computed from its contents and from the fingerprint of the row before, so changing one old row would break every later one.
  • No selection. Every published prediction counts, including the bad weeks. A game that is called off makes its row void. It is not quietly removed.

A track record starts short, and that is uncomfortable. A model with thirty published predictions has almost nothing to show. The temptation is to pad the number with the backtest. EdgeBoard does not. The two are labeled “Published before the game” and “Backtest (computed after the fact)”, scored separately, and shown side by side so you can compare them without ever mistaking one for the other.

Four questions for any “proven” model

  1. Were these numbers public before the games? If not, it is a backtest, whatever it is called.
  2. Which seasons were the settings tuned on? Results on those seasons are the least trustworthy.
  3. Is every prediction in the table? Ask about the ones that are missing.
  4. What did a simple baseline score on the same games? Beating a coin flip is not much. Beating “always pick the home team” is the first real hurdle.

If a model’s owner answers all four without flinching, the record is worth reading. You can put the same questions to EdgeBoard. The answers are on the Record page and the Method page.