Skip to content

Module 05

Measurement

Choose three measurements you can defend, including against how you would game them.

Motion is easy to produce and easy to mistake for progress.

Movement IIICheck whether it held. Choose measurements that resist gaming, then run the baseline again under the new manual and write an honest verdict — including that nothing improved.

Why the usual counts stopped working

Commits, lines, pull requests and tickets closed were always weak proxies, tolerated because producing them took human effort — the effort was the thing that made them mean anything. Agents remove that constraint, so the proxies now move whether or not anything improved, and they move most for the people delegating most indiscriminately.

A measurement that rises when you do the wrong thing faster is worse than no measurement, because it will be quoted.

The gaming test

For each candidate measurement, write down how you would inflate it if you were trying to. If you cannot think of a way, you have not thought hard enough. Then ask whether the inflating behaviour is one you would notice — and whether it looks different from real improvement.

The measurements that survive tend to share a shape: they count outcomes downstream of the work rather than the work itself, and they get worse when quality drops. Time from problem noticed to problem gone. Rework rate on accepted changes. How much of your week was repair. That third one you already have, from module 00.

Three is the limit on purpose. A dashboard of twelve is a way of never having to conclude anything.

Method

  1. 01List five candidate measurements.
  2. 02For each, write how you would inflate it deliberately.
  3. 03Discard any that can be inflated without the inflation being visible.
  4. 04Keep three. Write what each one would have to show for you to conclude the redesign failed.
  5. 05Record where each number comes from and how long it takes to collect. A measurement you will not gather is not one.

How this goes wrong

  • Measuring agent output

    Counting what the agent produced measures the tool's throughput, not your practice. The subject of this program is your working system.

  • No failure condition

    If you have not written what a bad result looks like, you will read every result as good. Define the disconfirming outcome before you collect anything.

  • Measurements you cannot actually gather

    An elegant metric requiring instrumentation you do not have will be quietly dropped in module 06, and its absence will be read as neutral.

The exercise

  • Choose three measurements, each with its gaming analysis written out.
  • For each, state the result that would mean the redesign did not work.
  • Note the collection cost. If it exceeds a few minutes a week, choose again.

Measurement set

The sections your artifact contains:

  1. Three measurements, and what each is a proxy for
  2. How each could be gamed, and whether that would be visible
  3. What each must show for me to conclude this failed
  4. Where the number comes from, and what it costs to collect

Done means. Each measurement has a written failure condition. Without one, module 06 cannot reach a verdict.

All modules

  1. 00Baseline
  2. 01Delegation Boundary
  3. 02Review Discipline
  4. 03Estate Legibility
  5. 04Blast Radius
  6. 05MeasurementYou are here
  7. 06Re-measurement