Top Ten Pitch
Ranked, argued, explained

F1 Rankings

Teammate comparison as the least-bad instrument in motorsport

Comparing two competitors in the same equipment is the closest motorsport gets to a controlled experiment, and its limitations are severe enough to require constant stating.

Teammate comparison as the least-bad instrument in motorsport
Teammate comparison as the least-bad instrument in motorsport · Photo via Pexels
Editorial note. Analysis and general information only — see our terms before acting on anything here.

Why the comparison is used

Two competitors in the same team operate equipment that is broadly similar, which removes the largest single source of difference in results. That makes the comparison far more informative than any figure drawn from finishing positions across the whole field. It is the standard instrument in motorsport assessment for exactly this reason, and it underlies most serious attempts at ranking.

The comparison is available in qualifying and race conditions, which allows different aspects of performance to be examined separately. Its value comes entirely from holding the machine approximately constant, and everything that weakens that assumption weakens the conclusion. That dependence is worth restating whenever the method is used, because the assumption is approximate rather than guaranteed in every situation.

The reference problem

A comparison tells you how somebody performed relative to one other competitor, whose own level is exactly what is unknown. Comfortably outperforming a weak reference is a much smaller achievement than narrowly outperforming a strong one, and the raw comparison cannot distinguish them. Placing the reference on a scale requires comparing them against their own other teammates, which extends the chain further.

Each additional link introduces uncertainty, and the total uncertainty grows faster than the number of links suggests. Conclusions about competitors separated by many links should therefore be held very loosely regardless of how precise the output looks.

Where equipment stops being equal

Development direction responds to feedback, and a team may pursue characteristics that suit one competitor more than the other. Component allocation, upgrade timing and reliability incidents are not always distributed evenly between two sides of a garage. Setup freedom means the machines can differ in configuration even when the specification is identical, which is a real source of variation.

Strategy calls differ between competitors within a race, and those calls affect finishing position independently of driving. None of these effects is usually large, and together they are comparable in size to the differences being measured.

Sample size within a season

A season provides a limited number of comparable sessions, and incidents, failures and interruptions remove several of them from the analysis. The remaining sample is small enough that a modest run of favourable circumstances can reverse the apparent order between two competitors. Qualifying comparisons offer more sessions and describe a narrower skill, since they exclude everything about managing a full race distance.

Race comparisons are more comprehensive and noisier, because a single early incident can determine the whole result. Reporting both, with the number of clean comparisons noted, is the minimum needed for a reader to calibrate.

Using the instrument honestly

The comparison supports statements about two specific competitors in a specific period and supports much less than that about careers. Aggregating many such comparisons into a network rating is a reasonable ambition, provided the resulting uncertainty is reported alongside. Where the network is sparse, the sensible output is a grouping rather than an ordering, since the evidence cannot separate close cases.

This is the same conclusion that every weakly connected ranking problem reaches, and motorsport is a particularly clear example. The instrument is genuinely the best available, and being the best available is not the same as being sufficient. Confusing the two is how a reasonable method becomes the basis for conclusions far more confident than the evidence behind it allows.

The short version
  • Same-team comparison controls the largest confounder
  • The reference competitor determines the result
  • Chained comparisons accumulate uncertainty quickly
F1 Rankingscomparisonmotorsportmethod
David Smith
Contributing writer, Top Ten Pitch

David Smith writes on f1 rankings for Top Ten Pitch, focusing on what the evidence supports rather than what makes the better headline.

Also by David Smith