Wimbledon
Head-to-head records as evidence, and where the sample runs out
A direct meeting between two competitors feels like the strongest possible evidence, yet the number of meetings is almost always far too small to support the conclusions drawn.

The intuitive appeal of direct meetings
A record between two competitors appears to settle the matter, because it seems to remove every confounding factor by placing them in the same contest. The intuition is sound in principle and weak in practice, since the number of such contests is rarely large enough to distinguish real differences. A handful of meetings can easily produce a lopsided record between two competitors of genuinely equal ability, purely through ordinary variation.
The record still feels decisive because it is concrete and memorable, which is exactly the property that makes small samples persuasive. Treating it as strong evidence therefore imports a great deal of confidence that the underlying data cannot support.
Why the meetings are not comparable
The contests in a head-to-head record are spread across years, surfaces, stages and conditions, so they are not repeated trials of the same question. One meeting may have occurred when a competitor was carrying an injury and another when the opponent was at an unusual low point in a season. Aggregating them assumes the circumstances were equivalent, which is precisely the assumption the spread of dates makes untenable.
A record accumulated mostly in one set of conditions tells you about those conditions rather than about the general relationship between the two. Splitting the record by context is more honest and usually leaves samples too small to interpret at all.
Selection effects in who meets whom
Two competitors meet only when both have advanced to the same round, which means the record is conditioned on both having played well enough to arrive. That condition is not neutral, since it filters out the occasions when one of them was performing poorly and exited earlier. A player who reaches late rounds consistently will therefore meet a rival mostly on the rival good days, which distorts the apparent balance.
The effect is strongest at the top of a draw and weakest in early rounds, so records built in different parts of the bracket mean different things. None of this is visible in the summary figure that gets quoted.
Where indirect evidence is stronger
A rating system built on the whole field pools thousands of results, and the resulting comparison between two competitors rests on far more information. It is indirect, which feels weaker, but the size and connectedness of the evidence base usually outweighs the directness of a small record. This is counterintuitive and it is the same reason a large well-designed comparison beats a small direct one in almost any measurement problem.
Direct records become genuinely informative when they are numerous and spread across contexts, which happens only for a few long rivalries. In those rare cases the record is worth taking seriously, and the rarity is the point.
Using the record without overreading it
A head-to-head can legitimately raise a stylistic hypothesis, such as one game shape troubling another, which other evidence can then examine. It cannot by itself establish that one competitor is better, and presenting it that way overstates what a short sequence of matches can show. Reporting the number of meetings alongside the record is a minimal courtesy that lets a reader calibrate immediately.
Noting the spread of surfaces and stages costs one further sentence and prevents most of the misreading that follows. Those two habits turn a misleading statistic into a modest and genuinely useful one.
- Head-to-head samples are small and unevenly distributed
- Context differences make meetings non-comparable
- Indirect evidence is often stronger than direct evidence





