Historical evaluation
How would our model have performed using only the information available before each historical match?
Model calibration
La Liga · 2020-21 → 2024-25 · 1,900 matches evaluated
When the model assigns a probability, how often does the outcome actually occur?
The closer to the diagonal, the better calibrated the model is.
Each match contributes three probabilities (home, draw, away). Ranges with fewer than 30 points are omitted.
By range
The most informative ranges and those that stray furthest from the diagonal.
- 15–20%model 17.6%observed 15.3%
- 20–25%model 22.9%observed 21.6%
- 25–30%model 27.3%observed 28.1%
- 30–35%model 32.2%observed 30.6%
- 50–55%model 52.3%observed 55.0%
- 65–70%model 67.2%observed 74.8%
- 80–85%model 81.8%observed 71.4%
Equivalent on the 0–2 scale: 0.592
How to read it
If the model assigns a probability close to 60%, we expect that outcome to happen roughly 6 out of 10 times across many observations.
Good calibration does not mean being right every time; it means the probabilities resemble the real frequency.
Historical evaluation, not hindsight.
Each simulated prediction uses only information available before the match.
Learn about our methodology