Club Rankings
Elo for clubs: rating a squad that is never the same twice
Pairwise rating systems were designed for individuals whose ability changes slowly, and applying them to clubs requires accepting that the rated entity keeps being replaced.

How the mechanism works
A pairwise rating adjusts two competitors after every meeting, transferring points from the loser to the winner in proportion to how surprising the result was. A predictable outcome moves the numbers very little, while an unexpected one moves them substantially, which makes the system self-correcting over time. Because the total is conserved, the ratings describe relative rather than absolute strength, and the whole scale can drift without meaning anything.
The design is elegant precisely because it requires only the result and the two prior ratings, with no external information at all. That minimalism is also its limitation, since it cannot use anything the result does not directly reveal.
The responsiveness parameter
The size of each adjustment is governed by a single parameter that determines how strongly one result influences the rating going forward. A large value tracks change quickly and produces a rating that jumps around after any unusual afternoon, which is unhelpful for seeding. A small value produces a smooth rating that lags behind genuine change, which is unhelpful when a club has been rebuilt.
Practitioners often vary the parameter by match importance, which is a reasonable heuristic and adds another choice requiring justification. Every published club rating differs partly because its designers chose this parameter differently, not because they observed different results.
The continuity problem
The system assumes it is rating a persistent entity whose ability evolves gradually, which describes an individual competitor reasonably well. A club is a different kind of object, because a substantial part of the squad can change during a single transfer window. The rating carried into a new season was earned by a group that may no longer exist in the same form.
Some implementations respond by pulling ratings towards the mean between seasons, which is an admission that the continuity assumption has failed. How far to pull is another free parameter, and it has a large effect on how quickly a rebuilt club is recognised.
Home advantage and other corrections
Venue effects are large enough that ignoring them biases the rating of any club with an unusual fixture pattern within a rating period. The standard fix adds a constant to the home side before computing the expected result, which is simple and assumes the effect is universal. It is not universal, since travel distance, altitude, surface and crowd conditions vary enormously between venues within the same competition.
Estimating a venue-specific effect requires more data than most clubs generate, so the constant persists as a workable approximation. The residual error falls unevenly, which means some clubs are systematically mismeasured by an amount nobody bothers to report.
What such a rating is good for
Pairwise ratings predict future results better than league position does, because they use margin, opponent quality and recency that a table discards. They are poor as a record of achievement, since they can rate a club highly at a moment when it has won nothing. Using them to seed competitions is defensible and unpopular, because participants prefer criteria based on results they earned rather than on a model.
That preference is not irrational, since a model can be adjusted and a record of trophies cannot. The sensible division is to use ratings for analysis and records for entitlement, which keeps each instrument within its competence.
- Elo rates a continuing identity, not a fixed squad
- The K factor sets responsiveness explicitly
- Transfer windows break the continuity assumption




