Running a Measurement Pilot
Testing a measure on one team before rolling it out finds the problems while they are still cheap to fix.
Measures are usually deployed everywhere at once and discovered to be wrong afterwards. A pilot is a quarter of work that prevents years of a bad number.
The measurement warning in “Running a Measurement Pilot” matters whenever software records work patterns. Organisations researching download time tracking software can use visit the source for time and project context, while outcomes, quality checks and direct feedback remain necessary to explain what the metric cannot show.
What a pilot is for
Finding out whether the measure can be collected consistently.
For an independent perspective related to “Running a Measurement Pilot”, consult the NIST Privacy Framework; it offers a useful external check on definitions, governance and the assumptions built into a proposed measure.
Finding out whether it means what you think.
Finding out how people respond to it.
And finding out what it misses, which the team will tell you within weeks.
Choosing the team
One that is representative rather than willing.
A cooperative team will make a bad measure work, which teaches you nothing.
And include the awkward case: the team with unusual work, because that is where definitions break.
What to specify beforehand
The definition, written: what counts as one unit, when it is counted, what is excluded.
The counterweight.
How long the pilot runs — a quarter, not a month.
And what would make you abandon the measure, which is the criterion nobody sets and the one that matters.
What to watch during
Whether two people recording the same thing produce the same number.
Whether the definition needed clarifying, and how often.
Whether behaviour changed, and how.
And whether anybody can explain what the number means without a document.
The consistency test
Give two people the same ten items and ask them to record them.
If the results differ, the definition is not operational and the measure will not be comparable across teams.
This takes an hour and catches the commonest failure, which is a measure that means different things in different places.
Asking the team
At the end: does this describe your work, what does it miss, and how would you make it go up without working better.
The last question produces the gaming list in advance, which is worth more than anything else the pilot generates.
The decision at the end
Adopt, adjust, or abandon.
Abandon is a successful outcome and is rarely taken, because the effort feels wasted.
It is not: a quarter spent not adopting a bad measure is a quarter well spent.
Rolling out
With the definition, the exclusions, the counterweight and the stated purpose.
And with what was learned in the pilot, including what the measure does not capture, which sets expectations correctly from the start.
What to check
Was your current measure piloted?
Would two people recording the same work produce the same number?
Was there an abandon criterion?
And did anybody ask how to inflate it?