Track record
Every prediction we make is graded against the real result — walk-forward, using only what was known before the match. No hindsight, no cherry-picking, and no "value bet" claims. This is the whole record.
592.457 matches · 9.729 players tracked across 9 leagues in 3 sports.
Monthly Brier from each match's pre-match prediction (walk-forward, no hindsight). A flat line below the 0.25 coin-flip mark means the model stays calibrated month after month — no drift.
Calibration
The real test of an honest probability: when we say 60%, does it happen about 60% of the time? Each dot is a band of predictions (543.964 graded); dot size = sample. On the dashed diagonal = perfectly calibrated — above it we were too cautious, below it too confident.
On the diagonal = honest. Above it = the model was too cautious, below = overconfident.
Our win probability vs the closing line
The hardest test isn't just against results — it's against the sharpest forecaster money can buy. For 21.098 matches we line up our stated win probability against the closing line — a top bookmaker's fair odds with the margin stripped out, the number the whole betting market converges to by tip-off. On average we sit -0.85 pp from it — no systematic over- or under-confidence — and 48.9% of our calls land within 5 pp of the closing line. We don't beat it — nobody reliably does — but this shows how close an honest model stays to the sharpest price. Source: Betradar closing line.
On the diagonal = honest. Above it = the model was too cautious, below = overconfident.
Sitting close to the closing line proves nothing on its own — you can hug the market and still be wrong. So here is the harder question, on the 20.985 of those matches that have since finished: who called them better? The honest answer is not us.
Same matches, same moment, scored the same way. The market is a few points ahead of us — it aggregates money, injuries, insider knowledge and last-minute news that a rating model never sees. We publish the gap instead of hiding it, because a forecast you can't check is worth nothing.
Set model vs market
Our set-win model — the same engine behind correct-score and handicap predictions — put to an independent cross-check. For 8.008 matches we compare our implied per-set win probability against the market's (a bookmaker's per-set odds, margin removed). On average we sit -0.09 pp from the market (typical gap 3.2 pp) — a check that our set distribution is market-consistent, not just fit to past results. Early sample; grows over time.
On the diagonal = honest. Above it = the model was too cautious, below = overconfident.
By sport & league
Table Tennis
Badminton
Padel
Totals & handicaps
Beyond who wins: how the model's points-total (over/under), set-handicap and points-handicap calls hold up against the real pre-game lines — same walk-forward grading, across table tennis & badminton.
Same honesty test for the handicap markets: our stated cover probability (favourite −1.5) vs how often it really covered. On the diagonal = honest.
Cover c→r = how often we called the favourite to cover vs how often they really did. The model slightly under-rates favourites' dominance — they cover a touch more than predicted.