Settled evidence

How model performance is measured

A useful evaluation names the metric, denominator, population, and timing rule. The current summary below is read from settled stored forecasts; when that evidence cannot be loaded, the site shows no substitute figure.

Current settled summary

235 scored math/poisson forecasts

Outcome accuracy

38.3%

Exact score accuracy

11.1%

Average 1X2 Brier

0.6511

235 probability records

Timing rule: only scored math/poisson forecasts stored before the match's recorded kickoff are included. Prediction dates run from 17 Feb 2026 to 10 Aug 2026 UTC.

Sample size reduces random movement but does not make different competitions or time periods directly interchangeable.

The evaluation record

Forecasts are stored before kickoff and joined to the final result after settlement. The results browser displays the recorded score pick beside the final score. Today's code is not rerun to replace yesterday's forecast, so the visible record remains auditable.

Outcome accuracy

Outcome accuracy is the share of scored forecasts whose leading home, draw, or away category matched the final outcome. It ignores the exact goal totals. The classes are not equally common, and rates vary by competition, so a percentage needs both its sample size and scope.

Exact score accuracy

Exact-score accuracy requires both predicted goal totals to equal the final score. It is much stricter than outcome accuracy: a predicted 2-1 and a final 1-0 share a home-win outcome but are not an exact match. Low exact-score rates are expected because probability is spread across many plausible cells of the score matrix.

Brier score

The 1X2 Brier score evaluates the full probability vector. For the actual outcome it compares the assigned probability with one, and for the other two outcomes with zero, then sums the squared errors. Lower is better. It rewards a model that assigns useful probabilities even when its largest category loses, and penalizes misplaced certainty.

Brier values should be compared only across the same outcome encoding and evaluation population. The live card reports a separate Brier denominator because legacy or non-probabilistic rows may not contain a score.

Calibration

Calibration groups comparable predictions by assigned probability and asks whether the event occurs at a similar observed rate. If home wins assigned around 60% happen about six times in ten over a large sample, that region is well calibrated. Calibration is not proven by one correct prediction and becomes unstable in narrow, small bins.

Confidence is not historical accuracy

The public confidence score summarizes model conditions, input quality, and residual uncertainty on a 0–100 scale. It does not mean that forecasts at confidence 70 have succeeded 70% of the time. The outcome probability is the model's event estimate; calibration and accuracy are observed after many matches settle.

Population limits

The current evidence scan combines supported competitions and historical settled forecasts. It is a broad track-record summary, not a controlled comparison of one competition or calculation change. A like-for-like study must define its population, date range, and sample size before comparing results.

Inspect the underlying evidence

Open the published results browser to compare match-level forecasts and final scores. Read the forecast methodology for the model pipeline and its limitations.