Skip to content
Both Halves Are Wrong

All notes / Traps

Comparing Teams That Are Not Comparable

Ranking teams by productivity is the most requested comparison and the one least likely to be valid.

Traps · Analysis

Two teams are measured, their figures differ, and somebody concludes one is better. Five things have to match before that conclusion is available, and they rarely do.

The practical lesson in “Comparing Teams That Are Not Comparable” is to connect every number to a decision and retain the context behind it. Teams exploring employment of relatives policy can review discover the official solution as one source of operational evidence, provided the purpose is disclosed and the interpretation is tested with the people affected.

The five

The same work: comparable items of comparable difficulty.

For an independent perspective related to “Comparing Teams That Are Not Comparable”, consult the Atlassian Team Playbook; it offers a useful external check on definitions, governance and the assumptions built into a proposed measure.

The same definition: what counts as one unit, what counts as done.

The same denominator: hours, heads, full-time equivalents, calculated the same way.

The same conditions: tooling, systems, demand pattern, customer mix.

And the same maturity: a team with three new starters is not the same team as one with none.

What usually differs

Work mix, which is the largest and least visible difference.

One team handles the escalations, the complex accounts, the awkward region.

They do the harder work and rank lower, which is both predictable and consistently missed.

The definition problem

Teams record things differently, particularly what counts as one item and when it is closed.

Two teams with identical performance can differ by a factor of two on recording convention alone.

Check the definitions before comparing anything, which takes a conversation and resolves a surprising share of apparent differences.

What a valid comparison looks like

The same team against itself over time.

Always valid, because the population and the work are the same.

And it answers the question most people actually have, which is whether things are getting better or worse.

When cross-team comparison is legitimate

Genuinely identical work: the same process, same items, different sites.

Here differences are informative and worth investigating — frequently they reveal a better local practice worth spreading.

That is the productive use, and it requires the five conditions to hold.

How to respond to the request

Do not refuse; reframe.

"Here is each team against its own trend, and here is what differs between them that would explain the gap."

That answers the underlying concern without producing a ranking that will be used, which the people section covers.

The league table problem

A ranking, once produced, is read top to bottom and the bottom is acted upon.

Whatever caveats accompany it.

Which is why the caveats must travel in the same table rather than in a footnote — and why, where the comparison is invalid, the right output is no table.

What to check

Do your teams use the same definitions?

How different is their work mix?

Has a ranking been produced, and what happened to the bottom?

And does anybody track each team against its own trend?