Va. The Gamed Measure
The oldest objection to any system that rewards measured performance has a name, or rather several. Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. Campbell’s law: the more a quantitative indicator is used for social decision-making, the more it distorts the processes it was meant to monitor.
The evidence is overwhelming, and it should be taken seriously.
How measures eat what they measure
Theodore Porter, the historian of quantification, tells the story of the US Forest Service, which raised its projected timber growth rates to “draw the teeth” from a law requiring sustainable harvests. When managers are judged by the accounts, Porter observes, they learn to optimise the accounts. Wendy Espeland and Michael Sauder showed how law school rankings remade the schools they ranked. Institutions reorganised themselves around the formula, not around legal education. Sally Engle Merry documented how global indicators crystallise over decades, locking in the choices of whoever designed them first.
The health evidence from the previous chapter adds a sharper finding. When the World Bank studied performance-based financing across dozens of countries, it found that paying for measured services increased idle capacity on the unmeasured ones. Providers did what they were paid for and dropped what they were not.
And the AI evidence adds a frightening one. Pan and colleagues found that as optimisers become more capable, they exploit misspecified objectives more aggressively, and that behaviour can shift abruptly at capability thresholds, with little warning. An organisation full of capable agents, each optimising a proxy, is Goodhart’s law with a turbocharger.
Why this does not defeat Axiocracy
Here is the uncomfortable truth that critics of measurement rarely confront: every organisation already runs on measures. Revenue, headcount, hours billed, papers published, tickets closed, visibility to the boss. The choice is not between measurement and no measurement. It is between measures that are chosen openly and revised deliberately, and measures that are chosen by default and never examined at all.
The worst gaming happens under three conditions: a single measure, chosen by someone else, frozen in place. Axiocracy’s design attacks each of them.
Plural measures, never one number. Value is a dashboard, not a scalar. A goal is expressed through several measures, weighted by agreement, with floors on the dimensions that must not be traded away (safety, ethics, quality). You cannot game a dashboard as cheaply as you can game a number.
Reward every valued dimension, or none. The idle-capacity finding is precise: neglect follows the boundary of what is rewarded. So the boundary must be drawn deliberately. If a dimension of the work matters, it goes into the measures. If it cannot be measured, it is protected outside the ledger rather than silently starved inside it.
Measure the increment, net of luck. Following the Cash on Delivery model, credit the value added above a baseline, not the raw output that would have happened anyway. Social return on investment practice supplies the deductions: deadweight (what would have happened regardless), attribution (what others contributed), displacement and drop-off.
Revise on a schedule. Measures get sunset clauses. Each cycle, participants review whether the proxies still track the goal, and change them if not. Divergence between the proxy and the real outcome is monitored like any other operational risk, because with capable agents it can arrive suddenly.
Keep islands of judgement. Numbers do not rule alone. Qualitative, deliberative review sits alongside the metrics, and the way qualitative judgement is translated into credit is itself documented.
Goodhart’s law is not an argument against Axiocracy. It is an argument against the unaccountable, single-number metric regimes we already live under. Axiocracy is the only proposal that makes the choice of measure a public, revisable, collective act, and that is the only known cure.