FIFA
The confederation problem: comparing teams that rarely meet
Rating systems depend on competitors playing each other, and international football is organised in a way that starves those systems of the connecting evidence they need.

Why connection matters more than volume
A rating system learns relative strength from matches, but it only learns about two groups relative to each other when members of those groups actually meet. A region can play an enormous number of internal fixtures and still tell the system nothing about how it compares with a region elsewhere. Internally, such a region will be finely ordered, because the dense internal schedule supplies exactly the comparisons the method requires.
Externally, its whole level floats, anchored only by the handful of matches its members play against outsiders in any cycle. The result is a table that is precise within regions and considerably less trustworthy across them, despite presenting one continuous list.
How isolated islands drift
When a group is weakly connected to the rest, its internal results circulate points among its own members without changing the level of the group. Repeated internal competition can inflate or deflate the whole cluster relative to reality, and nothing inside the cluster can detect the error. Qualifying formats that keep regions apart until a late stage extend this drift across years rather than weeks.
The correction, when it comes, is concentrated into a small number of high-profile fixtures, which makes it look like a shock rather than an adjustment. That is why the periodic surprise of a supposedly weaker side competing well is often a rating problem rather than a performance one.
Weighting attempts and their limits
Some systems apply regional strength coefficients so that a win inside a stronger region is worth more than the same win elsewhere. Those coefficients have to come from somewhere, and the only honest source is the same sparse set of cross-regional fixtures already identified as thin. A coefficient updated infrequently becomes a fossil, encoding a balance of power that may have changed considerably since it was last calculated.
Updating it frequently makes it noisy, since each cycle offers too few matches to estimate a regional level with any stability. The choice between a stale coefficient and a jumpy one has no comfortable resolution, which is why designs differ so widely.
What tournaments actually reveal
A global tournament is the rare occasion when many cross-regional fixtures happen in a short window, which makes it unusually informative for rating purposes. Even then the sample is small, the fixtures are not randomly assigned and the conditions favour some participants for reasons unrelated to strength. Knockout formats compound the problem, because the sides that would supply the most useful comparisons are eliminated before they can supply them.
A single tournament therefore adjusts regional estimates modestly rather than resolving them, and several cycles are needed before the picture settles. Reading one tournament as a verdict on regional strength repeats exactly the error the rating system is trying to avoid.
Reading a global table with appropriate caution
The practical advice is to trust the ordering within a region considerably more than the ordering across regions, especially in the middle of the table. Differences of a few places between sides from different confederations usually sit well inside the uncertainty the connection structure permits. Larger gaps are more meaningful, but even they should be treated as ranges rather than as points on a precise scale.
Seeding decisions built on the table inherit this uncertainty, which is one reason draws sometimes produce groups that nobody considers balanced. Naming the limitation openly would improve public understanding far more than another adjustment to the formula.
- Ratings need cross-group fixtures to be comparable
- Regional schedules create isolated rating islands
- Tournament results correct the drift only slowly





