1. Freeze predictions before the event
Store the event, market, probability, price snapshot, publication time and model version before the event starts. Editing a record after the outcome creates hindsight bias.
Keep wins, losses, refunds and excluded events visible. The exclusion rule should be written before results are known.
2. Separate training from future evaluation
Sports data are ordered in time. A realistic evaluation trains or tunes on earlier information and measures later events that were not used to choose the model.
Randomly mixing future and past observations can leak information. Rolling or walk-forward evaluation better resembles repeated real-world use.
3. Compare several complementary metrics
Accuracy answers whether the most likely label was correct. Brier score measures squared probability error, while log loss penalizes highly confident errors especially strongly.
Calibration compares stated probabilities with observed frequencies. No single number describes discrimination, confidence, settlement and practical usefulness at once.
- Accuracy → correct labels
- Brier score → probability error
- Log loss → confidence-sensitive error
- Calibration → whether percentages match frequencies
4. Use meaningful baselines
Compare against simple historical frequencies, a no-skill base-rate model and, when appropriate, a consistently captured margin-adjusted market probability.
A complex model that cannot improve on a transparent baseline may add presentation without adding predictive information.
5. Report sample size and uncertainty
A short run can be dominated by chance, especially for rare outcomes or narrow market segments. Publish counts by sport, market and probability range.
Confidence intervals or bootstrap ranges can show how uncertain the observed difference remains. They are more honest than treating one sample estimate as a permanent property.
6. Monitor drift after launch
Team strength, competition structure, data feeds and market behaviour change. Track performance by time period and model version rather than pooling every historical prediction forever.
A documented retraining and versioning process helps distinguish genuine change from selective reporting.
7. Results and EV measure different layers
Prediction metrics evaluate probability quality. Expected value adds a quoted price and therefore also depends on timing, margin and settlement assumptions.
A transparent report should show both the forecast layer and the decision layer without turning either into a guarantee.