Historical evaluation

How would our model have performed using only the information available before each historical match?

Market

Model calibration

La Liga · 2020-21 → 2024-25 · 1,900 matches evaluated

When the model assigns a probability, how often does the outcome actually occur?

0%0%20%20%40%40%60%60%80%80%100%100%Estimated probabilityObserved frequency
ModelIdeal calibration

The closer to the diagonal, the better calibrated the model is.

Each match contributes three probabilities (home, draw, away). Ranges with fewer than 30 points are omitted.

By range

The most informative ranges and those that stray furthest from the diagonal.

  • 15–20%model 17.6%observed 15.3%
  • 20–25%model 22.9%observed 21.6%
  • 25–30%model 27.3%observed 28.1%
  • 30–35%model 32.2%observed 30.6%
  • 50–55%model 52.3%observed 55.0%
  • 65–70%model 67.2%observed 74.8%
  • 80–85%model 81.8%observed 71.4%
0.296
Brier ScoreMeasures how close the probabilities are to what happened. Lower is better (0 = perfect).
Lower is better

Equivalent on the 0–2 scale: 0.592

How to read it

If the model assigns a probability close to 60%, we expect that outcome to happen roughly 6 out of 10 times across many observations.

Good calibration does not mean being right every time; it means the probabilities resemble the real frequency.

Historical evaluation, not hindsight.

Each simulated prediction uses only information available before the match.

Learn about our methodology