Our First 40 Graded Forecasts: The Model Has a Draw Problem
Forty settled predictions, 42.5% correct. Eleven matches finished level and the model picked the draw in none of them — not once in forty. Here is the number and what we are doing about it.
Two days ago we started recording every forecast before kick-off and grading it automatically when the match ends. Forty are now settled. This is the first honest look at the scoreboard, including the part that doesn't flatter us.
17 correct out of 40. That's 42.5%.
The number behind the number
Eleven of those forty matches finished level. That's 27.5% of the sample, roughly what you'd expect from football.
The model picked the draw zero times. Not once in forty attempts.
So every one of those eleven draws was an automatic miss. They account for 11 of our 23 misses — just under half. Strip them out and the record on matches that produced a winner is 17 from 29, or 58.6%.
That gap is the whole story. The model reads matches well enough when someone wins. It does not have a working concept of the stalemate.
We said this two days ago, in the other direction
On 10 September we published a look at the Champions League opening round and reported that the market priced the draw higher than our model in all five fixtures we examined. We flagged it as the clearest signal in that dataset and said it pointed toward avoiding the draw.
Two of those five finished level. PSV were 65% at home and drew with Shakhtar. Roma were our 51% pick in Istanbul and drew with Fenerbahçe.
The market was right. We were wrong, and we were wrong in exactly the direction we had just told readers to lean. Publishing that correction is the point of keeping a ledger.
Where it cost the most
Heavy favourites dropping points is the expensive version of this blind spot:
| Match | Model's favourite | Result |
|---|---|---|
| AZ Alkmaar – Willem II | AZ 81% | draw |
| Chelsea – Hull City | Chelsea 76% | draw |
| Al-Qadisiyah – Ettifaq | Al-Qadisiyah 75% | draw |
| Liverpool – Fulham | Liverpool 65% | draw |
Four sides priced as near-certainties by our own numbers, four dropped points. A model that never assigns meaningful weight to the draw will keep walking into these.
The Premier League round was a wipeout
Five Premier League fixtures settled. We got none of them.
- Bournemouth – Brentford: we said home, it was a draw
- Crystal Palace – Ipswich: we said home, Ipswich won
- Aston Villa – Nottingham Forest: we said home, Forest won
- Chelsea – Hull City: we said home, draw
- Liverpool – Fulham: we said home, draw
Nought from five, and a pattern inside it: every single pick was the home side. Three of the five ended level. Early-season English football has been messier than our inputs expected, and home advantage is doing less work than the model assumes.
What happens now
Forty matches is not enough to redesign a model on, and we're not going to pretend otherwise. It is enough to name a specific, testable defect: the draw probability is too low, and the home-win probability in tight domestic fixtures is too high.
Two things follow. First, we are treating any pick where our draw probability sits under about 25% in an evenly matched fixture as suspect until the sample grows. Second, we'll publish this same scoreboard again at 150 settled forecasts, when a hit rate starts carrying a confidence interval worth quoting rather than a shrug.
Until then, read our percentages as what they are: a model's opinion with a known blind spot, published before the match and left standing afterwards whether it worked or not.
The running record, including every miss above, is on the track record page. How the numbers are produced is in the methodology.
18+. Nothing here is advice to bet. A model that is 42.5% accurate on its own picks is not a betting system.