Complete backtest record
Run on 2026-08-20 and not changed after that. These are the only accuracy figures we publish across the site.
The model was trained with 5 leagues from the 2023-2025 seasons — 5,256 played matches — then checked against the 2025/26 season, which it had not used before. That season contained 1,752 fixtures: 1,709 were scored, while 43 were left out because a newly promoted team did not yet have enough history for the form window to be useful. Leaving those matches out is the fair approach; pricing them and then hiding them from the total is not.
We show both tests exactly as they finished. The first test uses a naive baseline, meaning it ignores the teams and predicts each outcome from how often that result has happened before. The model beat that baseline in 5/5 leagues at statistical significance. The second test is calibration: when we label something as 60%, it should land close to 60% of the time. It passed in 4/5.
Calibration error by league
Mean gap between the probability we assigned to a 1X2 pick and how often that pick landed. Lower is better; our own tolerance is 0.05.
Each bar is scaled against the weakest league in this group, so the visual difference looks bigger than the real difference. In absolute terms, the largest gap here is 5.30 percentage points. The amber league is the one outside our own tolerance, and that is why it remains visible here.
Why “99% accurate” is not a claim we make
You will see phrases like “99% accurate”, “zero risk”, and claims about certain winners around football tips online. Take “99% accurate” literally. It means a model keeps finding outcomes that fail once in a hundred, in a sport where a red card in the eleventh minute is still a normal Saturday. If our model began showing numbers like that, the calibration test on this page would flag it as a bug, not a triumph.
A near-certainty claim deserves to be pulled apart, not just waved away. Look at the sample behind it, and whether that sample was picked before or after the results were known. Look at what “accurate” has been defined to mean, because that definition often carries the whole claim and is rarely printed beside the number. Also look at what would make the claim false. A number in an advert costs nothing to type, and a claim that cannot fail is the cheapest kind.
The standard to use on us, and on anyone else, is whether the test could have gone badly. Ours could: 1,709 predictions from a season the model had never seen, published with the misses included in the total and with the league that failed our calibration tolerance named on this page.
None of this makes a prediction advice. The probabilities here are estimates, and gambling can still lose money whatever the number says. If it is no longer fun, our responsible gambling page is the better place to read next. 18+.
What this page does not show, and why
Three things readers may fairly expect. They do not exist yet, so we do not draw them.
- A live hit rate for this site. A competition or club only shows its own settled record after it reaches 100 resolved predictions of its own. Before that, one weekend can move the figure by several points, so publishing it would make a coin toss look like a track record. If a competition or team page has no accuracy block, this threshold is the reason.
- A comparison with the closing market. For a pre-match model, the fair benchmark is the price at kick-off. Measuring that needs our own archive of snapshots taken at known times. We are still building that archive. Using another forecast in its place would answer a different question.
- A profit-and-loss curve. Returns from a staking plan reflect the staking plan as well as the model, and this is one of the easiest charts online to make look flattering. Calibration and a baseline comparison are tougher to dress up, so those are the checks we publish instead.