1. The binary Brier score is squared probability error
For one binary forecast, subtract the outcome — 1 if it happened, 0 if it did not — from the predicted probability and square the difference. Average those squared errors across the sample.
Lower is better. A forecast of 70% that occurs has an error of (0.70 − 1)² = 0.09. If it does not occur, the error is (0.70 − 0)² = 0.49.
- Forecast 50%, event occurs → 0.25
- Forecast 70%, event occurs → 0.09
- Forecast 90%, event fails → 0.81
2. Confidence changes the penalty
A wrong high-confidence probability receives a much larger penalty than a wrong cautious probability. This makes the score useful when a service publishes probabilities rather than only labels.
The metric therefore discourages unsupported certainty, but it should still be interpreted against the event type and a relevant baseline.
3. Calibration asks whether probabilities mean what they say
Group comparable predictions into probability ranges and compare the average forecast with the observed frequency. If events assigned around 70% occur around seven times in ten over a large relevant sample, that range is reasonably calibrated.
Small samples can fluctuate substantially. A calibration table should show the number of observations in every range rather than presenting percentages without context.
4. Calibration is not the same as usefulness
A model that always predicts the base rate may look calibrated while providing little separation between events. Useful evaluation also considers resolution: whether probabilities meaningfully differ when outcomes are more or less likely.
For market comparison, a model can also be measured against a transparent baseline such as normalized market probabilities captured at a consistent time.
5. Multi-outcome sports need a declared convention
Football 1X2 has three outcomes. An evaluator can sum the squared errors across home, draw and away probabilities, but the scaling convention must be stated so values can be compared correctly.
Totals and handicap outcomes may be treated as binary only when pushes, voids and settlement rules are handled consistently.
6. A credible report shows the whole sample
Publish the sample period, number of settled predictions, exclusions, scoring formula and results by probability range. Keep misses and voids visible.
Brier score, calibration, accuracy and market comparison answer different questions. Reporting them together is more informative than selecting whichever number looks best.