Upsets vs expected: how to read our tennis predictability pages
Our ATP and WTA predictability boards say things like "13 upsets where the model expected 7.6". This page explains exactly what that means, how to read the chart on each player profile, and the statistics that decide who gets called a Wild card.
The model expects to be wrong. The question is how often.
Every match we publish a win probability. A probability is not a pick: when we say 55/45, we are saying the favorite should lose 45 of 100 such matches. So the fair question for any player is never "did the favorite always win?" but "did upsets happen about as often as the probabilities said they should?"
Take a concrete case. Suppose a player's ten matches carried favorite probabilities of 90, 85, 80, 75, 70, 65, 60, 55, 55 and 52 percent. The chance of an upset in each is one minus those numbers: 10, 15, 20, 25, 30, 35, 40, 45, 45 and 48 percent. Add them up and the model expects about 3.1 upsets across those ten matches even if it is perfectly calibrated. If two upsets happen, that player ran slightly calmer than priced. If six happen, something about them keeps breaking the script.
That sum is the expected upsets number on the boards. It is schedule-adjusted by construction: a top seed cruising through early rounds carries a low expected count, a qualifier grinding through coin-flip matches carries a high one, so neither is punished for the draw they happened to face.
Reading the chart on a player page
Each player profile shows two lines across the season. The dashed blue line is the model's expected upset count, accumulating match by match. The red line counts the upsets that actually happened. Three patterns cover almost everything you will see:
- Red hugging blue. The season went to script. Whatever the player's ranking, the model read them well: this is the "On serve" zone, and most established players live here.
- Red flat under blue. Fewer upsets than priced. The player was steadier than even the model dared to say: favorites held, underdog days stayed rare. That is a "Straight sets" season.
- Red climbing away above blue. Script-breaking results piling up: favorites (them or their opponents) kept falling. The gap between the lines at season's end is exactly the "extra upsets" number on the board. That is a "Wild card" season, and the steps in the red line tell you when it happened.
The step shape matters: a cluster of steps in a single month is a story (an injury, a surface switch, a form collapse), while evenly spaced steps just mean a player whose matches are genuinely hard to call.
Why most players are "On serve", and why that is honest
Upsets are random by nature, so over 20 or 60 matches the actual count will wander around the expected one even when the model is exactly right. Before we call anyone a Wild card we ask whether their gap is bigger than that natural wobble for their sample size (roughly a 90 percent confidence test on the upset count). With a season of data, most players do not clear that bar, and the board says so instead of inventing drama. The two names highlighted at the top of each board are the season's extremes either way. When even they sit inside the noise band, we say that too.
The math behind the verdicts: the Brier delta
Counting upsets treats a 51/49 miss and a 90/10 miss the same, which is readable but slightly crude. Behind the scenes the verdicts are cross-checked with a proper scoring rule. Each prediction gets a Brier score, the squared gap between the stated probability and the outcome. A calibrated model knows its own expected Brier on every match: for a probability p it is p times (1 minus p), largest at 50/50 and tiny for heavy favorites. Averaging realized minus expected Brier over a player's season gives a threshold-free version of the same question, where a near-coin-flip upset moves the number barely at all and a 90/10 shock moves it a lot.
The two measures agree strongly on our data (correlation 0.82 across players and seasons), so the boards speak in upsets, the units people actually think in, while the scoring rule keeps the verdicts honest.
What this deliberately does not claim
We tested whether "wild card" behaviour persists: it barely does. A player's gap in one season tells you almost nothing about the next (correlation around 0.1 to 0.2 season over season), even though two independent versions of our model agree strongly about the same season. Being hard to predict is mostly a season condition, not a personality. That is why every verdict is stamped with its season, why sample sizes and intervals are always shown, and why you will never see a career "predictability rating" here.
One more honesty note: recent seasons are scored on our as-published daily record, the same reconciled predictions behind the performance pages. Earlier seasons are scored with a strict out-of-sample backtest of the current model, and each season uses one source only, never a mix. Our calibration is public on the model transparency page, which is what makes the "expected upsets" baseline meaningful in the first place.
Explore the boards: most predictable ATP players · most predictable WTA players.