Top
Why the same results produce different orders under different formulas
Two rating systems fed identical fixtures will disagree, and the disagreement is not a bug but a direct consequence of what each formula was built to reward.

One dataset, several defensible answers
Give two analysts the same season of fixtures and they will return different orders, without either of them having made an arithmetic error anywhere. The divergence comes from the questions their formulas were built to answer, since one may be estimating current strength and the other summarising the whole campaign. A system that treats every fixture as equally informative will rank differently from one that discounts the earliest matches as stale evidence about a changed side.
Neither approach is dishonest, and both can be defended clearly, which is exactly why publishing a single order without a method invites unresolvable disputes. The useful reaction to two tables that disagree is curiosity about their assumptions rather than a search for the one that is correct.
Margin of victory as a contested input
Whether the size of a win counts is one of the sharpest dividing lines between rating systems, and it reshapes the order more than almost any other choice. Counting margin makes a system more responsive, because a decisive result carries more information about underlying strength than a narrow one decided late. Ignoring margin protects the system against teams that run up scores in favourable circumstances and against sports where late-game situations distort the final difference.
Systems that split the difference apply a cap, so that beyond a certain point additional goals or points stop adding to the rating at all. That cap is a compromise rather than a solution, and where it sits changes which competitors gain most from their heaviest wins.
Opponent quality and the circularity it creates
Adjusting for the strength of the opposition is obviously desirable, yet it introduces a loop, because opponent strength is itself an output of the same system. Iterative methods resolve that loop by passing repeatedly over the fixture list until the ratings stop moving, which is elegant but sensitive to starting assumptions. In a competition where groups rarely play each other, the loop has too little connecting evidence, and the resulting adjustments rest on very thin bridges.
A handful of cross-group fixtures can therefore carry enormous weight in setting the relative level of two otherwise separate populations of competitors. That fragility explains why ratings in sports with limited interconnection swing more than the underlying performances plausibly justify.
Recency weighting and the shape of memory
Every system must decide how quickly old evidence stops counting, and that decay curve is one of the least discussed yet most consequential settings available. A steep curve produces a table that tracks form, which is genuinely useful for anyone trying to describe the immediate present of a competition. A shallow curve produces something closer to a seasonal summary, which is better suited to awards, seeding and any decision meant to reward a body of work.
Problems arise when a table built with one curve is used for a purpose that requires the other, which happens constantly in practice. The curve also determines how long an unusual result keeps distorting the order after everyone involved has stopped considering it relevant.
Treating divergence as evidence
When several independently built systems agree on a competitor, that agreement is meaningful, because the conclusion survived quite different assumptions about what matters. When they disagree sharply, the disagreement localises the argument, since it points at the specific input that the systems value differently. A competitor rated highly by margin-sensitive systems and modestly by result-only systems is telling a coherent story about how their wins are achieved.
Reading several tables side by side is therefore more informative than trusting one, even when one of them carries the most familiar name. The comparison converts a shouting match about order into a discussion about criteria, which is the only version of the argument that can progress.
- Identical data supports many defensible orders
- Margin, opponent quality and recency each pull the order differently
- Disagreement between systems is informative rather than embarrassing



