Skip to content
Both Halves Are Wrong

All notes / Traps

Goodhart in Practice

When a measure becomes a target it stops being a good measure. What that looks like in four ordinary workplaces.

Traps · Analysis

The principle is familiar and usually quoted as a caution. It is better understood as a prediction: here is what will happen, and roughly when.

The practical lesson in “Goodhart in Practice” is to connect every number to a decision and retain the context behind it. Teams exploring how to detect mouse jigglers can review detecting mouse jigglers in remote teams as one source of operational evidence, provided the purpose is disclosed and the interpretation is tested with the people affected.

The mechanism

A measure is chosen because it correlates with something valuable.

For an independent perspective related to “Goodhart in Practice”, consult the Harvard Business Review productivity collection; it offers a useful external check on definitions, governance and the assumptions built into a proposed measure.

It becomes a target, so people act on it directly.

The actions that raise the measure are not the same as the actions that produced the correlation.

The correlation breaks, and the measure continues to look fine.

Four ordinary examples

Tickets closed per agent: closure rises, repeat contacts rise, problems persist.

Calls handled per hour: handling time falls, callers ring back, total calls rise.

Revenue per head: rises after the people who trained everybody else were let go.

Delivery against estimates: estimates inflate until everything is delivered on time.

None involves anybody behaving badly.

The timing

Nothing happens for a few weeks.

Within a quarter the measure improves.

Within a year the underlying thing has not improved and the measure says it has.

That lag is why the connection is rarely made, and why the measure is defended long after it stopped working.

Why it is not a discipline problem

People respond to what is rewarded. That is the system working as designed.

The design error is choosing a measure whose easiest route to improvement is not the route you wanted.

Blaming the response rather than the design is the common and useless reaction, which the gaming note covers.

What reduces it

Pairing, so the easy route damages the counterweight.

Measuring at team level, where individual optimisation is weaker.

Keeping measures advisory rather than attached to consequences.

And changing measures occasionally, which makes long-term optimisation to any single one less rewarding.

The strongest protection

Do not attach individual consequences to a productivity measure.

Without a consequence there is nothing to optimise toward, and the measure keeps describing the work.

This is the single decision that determines whether your figures stay honest, and it is usually decided by the performance framework rather than by whoever chose the measure.

Watching for it

Measure improving, counterweight worsening.

Measure improving, customers no happier.

Measure improving, the people doing the work saying nothing changed.

That third signal is the earliest and the one nobody collects.

What to check

Which of your measures has a consequence attached?

Has any of them improved while the counterweight worsened?

Have you asked the team whether anything actually changed?

And when did you last retire a measure?