There is a line that gets attributed to Peter Drucker at almost every leadership offsite: "What gets measured gets managed." The Drucker Institute says he never said it. The earliest version comes from a 1956 paper by V. F. Ridgway, and Ridgway was warning against it, not endorsing it.
The version I keep coming back to is Eli Goldratt's, from The Haystack Syndrome.
Tell me how you measure me and I will tell you how I will behave. If you measure me in an illogical way, do not complain about illogical behavior.
Eliyahu M. Goldratt, The Haystack Syndrome, 1990That is the whole argument. You do not get the behavior you want. You get the behavior you measure.
Engineering has learned this lesson at least three times. We counted lines of code and got long code. We counted story points and got inflated estimates. We counted tickets closed and got tickets split into halves. Each time the number went up, and each time the thing we actually wanted did not.
We are now doing it a fourth time, with AI.
What does your board slide show?
Pick the numbers on your AI slide. Each one is an instruction to your engineers. Here is what it tells them to do.
The new lines of code
Look at how most companies measure their AI rollout. The board slide says "AI adoption: 78%." Below it, if you are lucky, is a Copilot acceptance rate. Below that, tokens consumed per developer. Some companies have gone further. Microsoft told managers in 2025 to factor AI usage into performance reviews. Coinbase gave engineers a week to onboard an AI assistant and let go of some who did not. By early 2026 there were reports of token-usage leaderboards at several large tech companies, and a survey found that 58% of companies now require AI tool use.
None of these numbers are wrong to collect. All of them are wrong to target. They measure activity, and activity is the easiest thing in the world to produce. AI usage is an input. It is not output, and it is not value. Higher spend is not more software.
Ask an engineer to raise their acceptance rate and they will accept more suggestions. Ask them to raise token usage and they will generate more. Neither instruction says anything about whether the code was right, whether anyone reviewed it, or whether it shipped. Gergely Orosz put it plainly when he wrote about the leaderboards: the number of tokens a developer generates can easily be gamed, and if it is measured, it will be.
What the activity metrics hide
We now have enough independent data to see what happens when adoption is the target.
Faros AI looked at telemetry from more than 10,000 developers in 2025. Teams with high AI adoption completed 21% more tasks and merged 98% more pull requests. The PRs were 154% larger. Review time rose 91%. Bugs per developer rose 9%. At the company level, there was no significant correlation between AI adoption and improvement on any delivery or quality KPI. Their 2026 follow-up across 22,000 developers was starker.
The metric went up. The outcomes went the other way.
Change under high AI adoption, 2025 to 2026, teams in the Faros AI telemetry set.
View as table
| Measure | Kind | Change |
|---|---|---|
| Acceptance rate | Targeted metric | +200% |
| PRs merged with no review | Outcome | +31% |
| Bugs per developer | Outcome | +54% |
| Incidents per PR | Outcome | +243% |
Read that chart again. The metric everyone was tracking tripled. The outcomes went the other way. That is Goodhart's law rendered as a bar.
Google's DORA research found the same shape. In 2024, a 25% increase in AI adoption was associated with a 7.2% drop in delivery stability. In 2025, with nearly 5,000 respondents, throughput had turned positive, but stability was still negatively related to adoption. DORA's summary line: AI does not fix a team, it amplifies what is already there.
GitClear analyzed 211 million changed lines across 2020 to 2024 and found copy-pasted code rising, refactored code falling by more than half, and duplicated blocks up eightfold in a single year. The code got longer. It did not get better.
And METR ran the one randomized trial we have. Experienced open-source maintainers using AI tools took 19% longer to complete real tasks, while believing they had been 20% faster. METR's own 2026 update suggests newer tools have likely reversed the sign, and the study is being redesigned, so treat the exact number with care. But the gap between perceived and measured productivity is the durable finding. If you ask people whether AI made them faster, they will say yes. That is a feeling, not an outcome.
To be fair to the usage number
Adoption metrics are not useless. During a rollout, utilization is the right first gauge. You cannot measure the impact of a tool nobody has opened. Frameworks from DX and GitHub both treat usage as a leading indicator that precedes outcome metrics, and I agree with that.
The failure is not collecting the number. The failure is stopping there, or worse, promoting it to a target. The moment "AI usage %" appears on a performance review, it stops describing your organization and starts shaping it.