81.5% across 19,340 predictions on 2,546 matches the model had never seen. Below is the method, the per-stage breakdown, and the calibration table — the part that cannot be faked.
Accuracy claims are easy to inflate. The usual trick is to score a model on matches it was trained on, which measures memory rather than prediction. So the split here is by date, not at random:
Trained on: T20 matches up to 2024
Tested on: 2025-2026 matches, never seen during training
Sample: 2,546 matches, 19,340 predictions
Reproduce it:
python backtest_comprehensive.py --holdout
A prediction counts as correct if the side the model favoured went on to win. Win probability is scored at fixed points in every innings, not cherry-picked moments.
Predicting a finished match is trivial, so the number that matters is how early the model is useful. It is already close to its ceiling by the sixth over:
| Point in the innings | Accuracy | Predictions |
|---|---|---|
| After Over 6 | 79.2% | 5,036 |
| After Over 10 | 82.1% | 4,926 |
| After Over 12 | 82.6% | 4,820 |
| After Over 15 | 82.2% | 4,558 |
This is the question a hit-rate cannot answer. A model that simply backs whichever side is obviously ahead will score well on direction and still be useless, because its numbers mean nothing in between. Calibration checks the numbers themselves: group every prediction by what the model said, then look at how often those situations really ended in a win.
| Model said | Predictions | Actually won | Gap |
|---|---|---|---|
| 0-10% | 4,529 | 3.8% | 1.2 |
| 10-20% | 1,564 | 15.3% | 0.3 |
| 20-30% | 1,188 | 25.4% | 0.4 |
| 30-40% | 1,074 | 34.5% | 0.5 |
| 40-50% | 1,175 | 43.9% | 1.1 |
| 50-60% | 1,122 | 50.4% | 4.6 |
| 60-70% | 1,350 | 62.1% | 2.9 |
| 70-80% | 1,568 | 72.5% | 2.5 |
| 80-90% | 1,708 | 81.1% | 3.9 |
| 90-100% | 4,062 | 96.1% | 1.1 |
The curve climbs monotonically and the worst band is off by 4.6 percentage points. That is the evidence behind the headline figure, and it is published here rather than described, because a claim without a table under it is only a claim.
Two honest limits, stated because they change how the number should be read:
The figures above come from historical ball-by-ball files. Live matches run on a different data feed with its own quirks, so live predictions are recorded at overs 6, 10 and 15 before the result is known and scored afterwards. That record is kept separate from this one on purpose — a backtest and a live record are different claims, and merging them would flatter both.