Top Ten Pitch
Ranked, argued, explained

Club Rankings

Cross-league comparison and the calibration problem

Placing clubs from different domestic competitions on one scale requires bridging evidence that only continental fixtures provide, and there is never enough of it.

Cross-league comparison and the calibration problem
Cross-league comparison and the calibration problem · Photo via Pexels
Editorial note. Analysis and general information only — see our terms before acting on anything here.

Islands with strong internal order

Each domestic competition generates a dense web of results that supports a precise internal ordering with very little ambiguity. What it cannot generate is any information about how the whole competition compares with another one played elsewhere. The internal precision is therefore misleading, because it invites the assumption that the scale is comparable across competitions when it is not.

Two clubs can hold identical positions in their own tables while being separated by a considerable margin in actual strength. Bridging that gap requires matches between the two populations, and those matches are rare, unevenly distributed and played under unusual conditions.

The bridge is narrow and biased

Continental competition supplies the connecting fixtures, but the clubs that take part are the strongest from each domestic competition rather than a representative sample. Inferring the level of a whole league from the performance of its leading clubs assumes the rest of the league scales similarly, which is unsafe. A competition can have an unusually strong top and an ordinary middle, and the bridge fixtures would report only the former.

The bias runs the other way as well, since a competitively balanced league may send clubs that are strong relative to their domestic rivals and modest overall. Any cross-league calibration therefore rests on an assumption about the shape of each distribution that the available data cannot check.

Timing and condition effects

Continental fixtures are played mid-season under compressed schedules, which affects clubs differently depending on their domestic calendar and squad depth. Travel, unfamiliar conditions and different officiating interpretations all contribute variation unrelated to the underlying quality being estimated. These effects are not random across leagues, since some calendars align far better with the continental schedule than others do.

A calibration built from such fixtures therefore contains a systematic component that looks like league strength and is actually scheduling. Correcting for it requires estimating each of those effects separately, which the small number of fixtures cannot support.

Transfer flows as indirect evidence

Movement of players between competitions offers a second, indirect source of information about relative level that does not depend on match results. Consistent flow in one direction suggests a difference in level, since competitors generally move towards greater opportunity and stronger competition. The signal is heavily contaminated by finance, regulation and geography, so flow direction reflects resources at least as much as playing standard.

It is nonetheless a genuinely independent piece of evidence, and agreement between flow and result-based calibration is more convincing than either alone. Disagreement between them is informative too, since it usually points at a competition whose resources and playing level have diverged.

How much precision to claim

The available evidence supports a rough grouping of competitions into bands rather than a precise ordering with meaningful gaps between neighbours. Presenting a continuous cross-league ranking implies a level of calibration the bridging fixtures cannot deliver, however sophisticated the model appears. Bands communicate the uncertainty honestly and still answer the question most readers are actually asking about relative standing.

Within a band, the sensible statement is that the competitions are not distinguishable on current evidence, which is a real finding rather than an evasion. Claiming less than the model output allows is usually the correct decision when the underlying connections are this thin.

The short version
  • Separate leagues form weakly connected rating islands
  • Continental matches are the only bridge and are few
  • Selection effects distort what those matches show
Club Rankingscalibrationcomparabilityclub ranking
David Smith
Contributing writer, Top Ten Pitch

David Smith writes on club rankings for Top Ten Pitch, focusing on what the evidence supports rather than what makes the better headline.

Also by David Smith