Skip to content
Both Halves Are Wrong

Measuring productivity

A ratio of two numbers you have not checked

Fifty notes for whoever has to report on productivity: what the ratio can and cannot support, what to measure when the work has no comparable unit, the traps that make a good-looking figure wrong, and why the answer is usually in the process rather than in the people.

Core notes remain method-focused; separate guides compare named tools. No unverified productivity-improvement figures. This collection argues for careful measurement, not against measurement.

  • Request
  • Waiting for approval
  • Work
  • Waiting for the specialist
  • Review
  • Waiting for the weekly run
  • Delivered

Measured end to end, most of the elapsed time in a typical process is queueing rather than working. Productivity measures see none of it.

Two approximations, divided

Productivity is output divided by input. Everything difficult about measuring it follows from the fact that both numbers are harder to establish than they look.

The measurement warning in “A ratio of two numbers you have not checked” matters whenever software records work patterns. Organisations researching employee monitoring software can use employee monitoring software for time and project context, while outcomes, quality checks and direct feedback remain necessary to explain what the metric cannot show.

The numerator counts whichever unit was easiest: tickets rather than problems solved, documents rather than decisions supported. It treats items of wildly different size as equal, ignores quality so that work done twice counts twice, and excludes anything without an obvious unit — which in most organisations is a large share of the work.

For an independent perspective related to “A ratio of two numbers you have not checked”, consult the OECD productivity resources; it offers a useful external check on definitions, governance and the assumptions built into a proposed measure.

The denominator is usually hours present, which include meetings, waiting, administration and interruption. The proportion that is actually work varies enormously between roles, teams and weeks.

Divide one by the other and two approximations become one number with decimal places. The approximation does not disappear; it becomes invisible. And the figure is then compared between teams and across quarters as though it were precise, which is where the decisions go wrong.

The remedy is one extra line in the reporting: show both halves alongside the ratio. "Output flat, hours down eight per cent after the vacancy freeze" contains everything the quotient contains and none of the false precision — and it prompts the right question.

Any measure that becomes a target

The principle is familiar as a caution. It is better understood as a prediction with a timescale.

Nothing happens for a few weeks. Within a quarter the measure improves. Within a year the underlying thing has not improved and the measure says it has. That lag is why the connection is rarely made, and why the measure is defended long after it stopped working.

Tickets closed per agent: closure rises, repeat contacts rise, problems persist. Calls handled per hour: handling time falls, callers ring back. Revenue per head: rises after the people who trained everybody else were let go. None of it involves anybody behaving badly — people respond to what is measured, which is the system working as designed.

The reliable protection is a counterweight: a second measure that falls when the first is inflated dishonestly. Volume with reopen rate. Output with rework. Throughput with lead time. One measure can be optimised; two that oppose each other require genuine improvement. And the counterweight has to be collected independently, because a quality score produced by the same people with the same incentive moves with whatever is being rewarded.

When gaming appears

The common reaction is to tighten enforcement. That addresses the symptom and discards the most useful signal available.

Gaming tells you three things: the measure does not match the work in the view of the people doing it, the easiest route to a better number is not the route you intended, and the consequence attached is large enough to be worth responding to. The person splitting work into smaller units knows the unit was never comparable. They are responding accurately to a measure that is inaccurate.

Ask instead what the measure is missing. The answer is usually specific and fixable, and the conversation takes an hour where enforcement takes years and does not work.

Work that has no unit

The arithmetic requires units to be alike. Two reports, one a summary and one a three-month investigation. Two tickets, a password reset and a corrupted database. Counting them equally is not an approximation; it is a different measurement entirely.

Inventing a unit does not help. Story points, complexity ratings and weighted case counts are estimates made by the people being measured, which means using them for assessment corrupts them immediately.

What works instead needs no comparable unit at all. Flow: how long things take from request to delivery. Rework: how often something comes back, which is unambiguous and measures waste directly. Waiting: how much of the elapsed time was queueing. And asking — structured and periodic — which is treated as soft and is consistently more useful than any ratio.

Sometimes the honest answer is that no reliable measure exists. A function that prevents incidents has no output to count and its measure improves when it fails. Saying so, with what you track instead, is a stronger position than a figure that collapses under examination.

Where the time actually goes

Measure a process end to end and most of the duration is usually waiting rather than working. This is consistently true and consistently unmeasured, because productivity measures count output and effort and waiting is neither.

An approval taking days where the decision takes minutes. A handoff where work sits in somebody's queue. A weekly meeting that gates everything. A specialist bottleneck nobody realised was one. The longest waits sit between two teams who both look busy, because neither team's own measures capture the gap.

Which makes waiting the highest-return thing to measure, by a considerable margin — removing a queue usually requires a decision rather than resources. A delegation threshold. A named deputy. A shorter batch interval. Most cost nothing and take effect immediately.

And one diagnostic question reorients most improvement efforts: if everybody worked ten per cent harder, how much more would be delivered? In a constrained system, almost nothing. It would queue.

Reading a distribution rather than an average

Productivity data is rarely symmetrical. Most items are quick and a few take a very long time, which produces a long tail in every duration measure and in most volume measures too.

The average sits above most of the data, moves when one extreme item appears, and describes nobody's experience — not the typical case, not the bad case. Report the median and a high percentile instead: most things take this long, and the slow ones take this long.

And the tail is the finding. The slow items are where the problems live: the exception handling, the escalation, the one waiting on a specialist. Improving the median is usually easy and worth little; improving the tail is where the customer experience is.

One habit prevents most of the explaining of random variation that fills management meetings: plot the series with its ordinary range marked. Anything inside is noise. Anything outside is worth a question.

Check the denominator first

When a productivity figure moves, attention goes to the output. The input is usually where the change actually happened.

A vacancy raises output per head immediately — reported as improvement, caused by being short-staffed, and it reverses when the post is filled, which then reads as a decline. A change in how full-time equivalents are counted shifts every figure in every team with no change in any work. A reorganisation moves people between teams. The annual headcount snapshot is taken on a different date.

So pull the headcount series alongside the output series before writing any explanation. Two minutes, and it resolves a large share of apparent productivity changes.

The same caution applies to sudden improvements generally: ask what left. People, work, customers or scope. Service measures look better when the hardest customers go elsewhere, and complaint rates fall because the complainers left.

The work that produces no visible activity

In professional work the part that determines the outcome frequently involves no observable action at all. Consider the two routes to a hard problem: think for twenty minutes and write the answer in five, or start typing immediately and iterate for two hours. The first is better work and registers as less of it.

The same inversion applies to expertise. An expert solves in ten minutes what takes somebody else two days, and every measure of effort counts that as less work. An expert measured on volume learns to take easier work, which is a catastrophic misallocation and entirely rational given the incentive.

And it applies to the work that enables others — the question answered, the review that caught a problem, the colleague taught, the documentation written. All of it appears in somebody else's output, if anywhere. Under individual measurement, helping declines: people close their doors, stop reviewing, stop teaching. Team output falls while every individual figure holds steady.

Most teams have one person everybody asks, whose individual output is lower by construction. Removing them on the basis of their figures is the single most damaging decision this kind of measurement produces, and it has happened in enough organisations to be a known pattern.

Measurement and trust

The same measure can improve an organisation or damage it, and the difference is not in the measure. It is what people believe it is for.

Measurement that is trusted produces accurate data, because nobody is protecting anything. It surfaces problems early, because raising one costs nothing. And it gets improved by the people measured, who suggest better measures than anybody outside the work would.

Measurement that is not trusted produces managed data, hides developing problems until they are unavoidable, and makes every future measurement initiative harder — including the ones that would have been fine.

One question tests which you have. Would somebody tell you about a developing problem before it showed in the figures? If the answer is no, the figures are describing something other than reality, however carefully they are collected.

What the core notes deliberately avoid

No product rankings appear inside the core notes. Named products are kept in separate comparison guides so the measurement analysis does not depend on one supplier.

No figures for average productivity improvement, because the studies producing them are commissioned by people selling something.

And no argument that measurement is bad. The argument is that bad measurement is worse than none, and that most of what goes wrong is avoidable with four or five habits that cost nothing.

0808sections

Reference

Terms, and where to start for the common situations.

Product comparisons

Tool guides for evidence-led teams

Three detailed shortlists covering productivity measurement, time tracking and workforce analytics, with pilot and governance checks.

7 Employee Productivity Measurement Tools That Avoid Vanity Metrics

Seven employee productivity tools compared by evidence quality, reporting, implementation effort and safeguards against misleading activity scores.

Compare 7 tools →

12 Time Tracking Tools for Reliable Project Reporting

Twelve time tracking tools compared for projects, timesheets, billing, automatic capture, corrections and responsible team adoption.

Compare 12 tools →

16 Workforce Analytics and Performance Platforms Compared

Sixteen workforce analytics and performance platforms compared for operational patterns, feedback, planning, reporting and responsible use.

Compare 16 tools →

The short version

Most productivity problems are not about people

They are queues, handoffs, approvals, old equipment and fragmented weeks. A per-person measure hides all of them and a per-item measure shows them, from data most organisations already hold.