1. Turning recent matches into a scoring rate
Each constant shown here is used in the code, not a rounded explanation.
Start with a club’s last 15 finished matches, ordered from newest to oldest, with the index i = 0, 1, 2 … The weight for each match is 0.9i. That makes the latest match count 1, the previous one 0.9, and the fifteenth about 0.23. From this weighted set we calculate an attacking figure and a defensive figure:
att = Σ wi·fi / Σ wi · def = Σ wi·ai / Σ wi · wi = 0.9i
Here f means goals for and a means goals against. Expected goals for and against are used instead when expected-goals records cover at least 60% of the window’s weight. Weight share matters more than match share: missing data from last weekend counts more than a missing record from December.
After that, both figures are pulled toward the competition’s own scoring average μ. That average is taken from its two most recent seasons, as they stood before this kick-off:
shrunk = (n·v + 5·μ) / (n + 5)
In this formula, n is the number of matches actually present in the window. Five is the prior strength. A club with fifteen matches keeps three quarters of its own number, while a club with three keeps a bit over a third. That prevents a promoted side’s two-match opening run from creating an overconfident rating.
2. Building the expected-goal rates
The shrunk numbers are combined into a rate for each team. They are measured against the competition’s home and away means (mh, ma), with μ as the midpoint:
λhome = mh · (atthome / μ) · (defaway / μ)
λaway = ma · (attaway / μ) · (defhome / μ)
Home advantage is not added afterwards. It is already built into mh and ma, based on that competition’s own results. If a per-league fit is available, λhome also includes a fitted home scaling factor. Both rates are then capped within 0.2 to 4.5 goals, so a bad input cannot make the model expect a 9-0 match.
3. The Dixon-Coles scoreline grid
Each team’s goals are first treated as independent Poisson counts. Then the four lowest-scoring outcomes are adjusted, because plain Poisson is measurably off for 0-0, 1-0, 0-1 and 1-1. That is the issue the Dixon-Coles correction is designed for. For scoreline i-j:
P(i, j) ∝ Pois(i; λhome) · Pois(j; λaway) · τ(i, j)
τ(0,0) = 1 − λhomeλawayρ
τ(0,1) = 1 + λhomeρ
τ(1,0) = 1 + λawayρ
τ(1,1) = 1 − ρ
τ(i,j) = 1 everywhere else
The grid covers 0 to 10 goals for each side, giving 121 cells, and it is normalised so the total equals one. ρ is fitted separately by competition where that fit exists. Where it does not exist, the fallback is −0.10, and the prediction is labelled with the fallback model version instead of the fitted version.
4. Taking markets from the grid
Each market shown on the site comes from summing the same 121 cells.
| Market | Grid definition |
|---|---|
| Home / Draw / Away | Σ cells with i > j · Σ cells with i = j · Σ cells with i < j |
| Over/under 2.5 | Σ cells with i + j ≥ 3; under is the remainder |
| Both teams to score | Σ cells with i ≥ 1 and j ≥ 1 |
| Correct score | a single cell |
| Double chance, draw, handicap, goals | the same cells added together in other groupings |
All these markets are different views of the same grid. So our over/under figure cannot conflict with our correct-score grid, and our double chance cannot conflict with our 1X2. When markets are priced separately, that consistency does not automatically follow. See each market, with its fixtures, here: prediction types.
5. Removing the margin and blending the price
The three selections are first turned into one price each. For every fixture, we use the median across bookmakers that quote the full three-way market. Those three prices imply probabilities whose total is more than one. The extra part is the margin, and removing it is a normalisation:
qk = (1 / pricek) / Σj (1 / pricej)
The published 1X2 probability is then a fixed-weight blend of that market view and our model view:
final = 0.25 · model + 0.75 · market
If no price is recorded, final stays as the model probability, and the prediction is marked as model-only. This step applies to 1X2 only — over/under, both teams to score and correct score are stored without any market component, so those numbers are the direct output of the grid above.
The difference between the two inputs is what the site calls an edge: model − market, in percentage points, on the same selection. It says the price may be wrong, not that the result is certain. The fixtures with the biggest gap are shown on value bets.
How this is measured
Equations alone do not prove performance. The testing is shown on the accuracy page: the hold-out season, the baseline comparison, the calibration error by league, including the league that missed our own tolerance.