POV
Measurement Engineering leadership

You get the behavior you measure. Not the one you want.

AI adoption is being measured the way lines of code were measured in 1995. The results are about the same.

The position

Stop targeting AI usage. Measure four team-level outcomes instead: lead time, failure and rework, cost per change, and promised versus landed. Every one of them can be computed from data you already have.

Skip the story, show me the four numbers

There is a line that gets attributed to Peter Drucker at almost every leadership offsite: "What gets measured gets managed." The Drucker Institute says he never said it. The earliest version comes from a 1956 paper by V. F. Ridgway, and Ridgway was warning against it, not endorsing it.

The version I keep coming back to is Eli Goldratt's, from The Haystack Syndrome.

Tell me how you measure me and I will tell you how I will behave. If you measure me in an illogical way, do not complain about illogical behavior.

Eliyahu M. Goldratt, The Haystack Syndrome, 1990

That is the whole argument. You do not get the behavior you want. You get the behavior you measure.

Engineering has learned this lesson at least three times. We counted lines of code and got long code. We counted story points and got inflated estimates. We counted tickets closed and got tickets split into halves. Each time the number went up, and each time the thing we actually wanted did not.

We are now doing it a fourth time, with AI.

1990s
Lines of code
Got long code
2000s
Story points
Got inflated estimates
2010s
Tickets closed
Got tickets split in half
2025
AI usage %
Got accepted suggestions nobody read
Try it

What does your board slide show?

Pick the numbers on your AI slide. Each one is an instruction to your engineers. Here is what it tells them to do.

The new lines of code

Look at how most companies measure their AI rollout. The board slide says "AI adoption: 78%." Below it, if you are lucky, is a Copilot acceptance rate. Below that, tokens consumed per developer. Some companies have gone further. Microsoft told managers in 2025 to factor AI usage into performance reviews. Coinbase gave engineers a week to onboard an AI assistant and let go of some who did not. By early 2026 there were reports of token-usage leaderboards at several large tech companies, and a survey found that 58% of companies now require AI tool use.

None of these numbers are wrong to collect. All of them are wrong to target. They measure activity, and activity is the easiest thing in the world to produce. AI usage is an input. It is not output, and it is not value. Higher spend is not more software.

Ask an engineer to raise their acceptance rate and they will accept more suggestions. Ask them to raise token usage and they will generate more. Neither instruction says anything about whether the code was right, whether anyone reviewed it, or whether it shipped. Gergely Orosz put it plainly when he wrote about the leaderboards: the number of tokens a developer generates can easily be gamed, and if it is measured, it will be.

What the activity metrics hide

We now have enough independent data to see what happens when adoption is the target.

Faros AI looked at telemetry from more than 10,000 developers in 2025. Teams with high AI adoption completed 21% more tasks and merged 98% more pull requests. The PRs were 154% larger. Review time rose 91%. Bugs per developer rose 9%. At the company level, there was no significant correlation between AI adoption and improvement on any delivery or quality KPI. Their 2026 follow-up across 22,000 developers was starker.

The metric went up. The outcomes went the other way.

Change under high AI adoption, 2025 to 2026, teams in the Faros AI telemetry set.

Percent change under high AI adoption Acceptance rate rose 200 percent. Over the same period, PRs merged with no review rose 31 percent, bugs per developer rose 54 percent, and incidents per PR rose 243 percent. 0 +100% +200% Acceptance rate +200% PRs merged, no review +31% Bugs per developer +54% Incidents per PR +243%
Source: Faros AI, "The Acceleration Whiplash", April 2026. 22,000 developers, 4,000+ teams.
View as table
MeasureKindChange
Acceptance rateTargeted metric+200%
PRs merged with no reviewOutcome+31%
Bugs per developerOutcome+54%
Incidents per PROutcome+243%

Read that chart again. The metric everyone was tracking tripled. The outcomes went the other way. That is Goodhart's law rendered as a bar.

Google's DORA research found the same shape. In 2024, a 25% increase in AI adoption was associated with a 7.2% drop in delivery stability. In 2025, with nearly 5,000 respondents, throughput had turned positive, but stability was still negatively related to adoption. DORA's summary line: AI does not fix a team, it amplifies what is already there.

GitClear analyzed 211 million changed lines across 2020 to 2024 and found copy-pasted code rising, refactored code falling by more than half, and duplicated blocks up eightfold in a single year. The code got longer. It did not get better.

And METR ran the one randomized trial we have. Experienced open-source maintainers using AI tools took 19% longer to complete real tasks, while believing they had been 20% faster. METR's own 2026 update suggests newer tools have likely reversed the sign, and the study is being redesigned, so treat the exact number with care. But the gap between perceived and measured productivity is the durable finding. If you ask people whether AI made them faster, they will say yes. That is a feeling, not an outcome.

To be fair to the usage number

Adoption metrics are not useless. During a rollout, utilization is the right first gauge. You cannot measure the impact of a tool nobody has opened. Frameworks from DX and GitHub both treat usage as a leading indicator that precedes outcome metrics, and I agree with that.

The failure, precisely

The failure is not collecting the number. The failure is stopping there, or worse, promoting it to a target. The moment "AI usage %" appears on a performance review, it stops describing your organization and starts shaping it.

The point of view

The behavior you actually want

Nobody wants adoption. Adoption is a proxy. What a CTO wants, if you make them finish the sentence, is roughly this: ship faster without breaking more, spend less per change, and land what we promised the business we would land.

Every clause in that sentence is measurable. None of them are measured by acceptance rate.

Four questions, four numbers, all at the team level
  1. Did we ship faster?
    Lead time for changes.
  2. Without breaking more?
    Change failure rate, rework rate, incidents per merged PR. If lead time improved 20% while change failure rate rose, would anyone notice before the board did?
  3. At a lower cost per change?
    Cost per merged, shipped change, across tokens, seats, pipelines, and cloud. Cost per developer stopped describing the spend the day developers started spending tokens.
  4. Did we land what we promised?
    The scope that was asked for, reconciled against what was actually built and verified, before the release date moves. Merged is never done. Verified is done.

Three properties matter more than the specific list.

  • Each one can only go up when the outcome does. You cannot inflate "incidents per PR" by pasting harder.
  • They are team-level by design. The lesson of lines of code was not only that the number was wrong. It was that ranking individuals on it corroded the team. Any AI metric that becomes an individual scorecard will produce the same corrosion faster, because the tools make the number so easy to move.
  • None of them require a vendor. Lead time, failure rate, rework, cost per change, and asked-versus-built can all be computed from the version control system, the pipeline, the ticket tracker, and the invoices you already have. If a metric cannot be audited from your own data, it is not a metric. It is a slide.

Measurement is a message

The dashboard is the loudest manager in the building. It never takes a day off, and every engineer knows exactly what it wants. If it wants tokens, it gets tokens. If it wants accepted suggestions, it gets accepted suggestions and a review queue nobody clears.

So the real work of an AI rollout is not choosing the tool. It is choosing the number. Pick the one that describes the outcome you would defend to your board, put it at the team level, and let usage take care of itself.

Engineers are not gaming your metrics. They are obeying them. Give them something worth obeying.

Sources
  1. Drucker Institute, "Measurement Myopia" (2013), on the misattribution; V. F. Ridgway, "Dysfunctional Consequences of Performance Measurements", Administrative Science Quarterly, 1956.
  2. Eliyahu M. Goldratt, The Haystack Syndrome, North River Press, 1990.
  3. Faros AI, "The AI Productivity Paradox" (July 2025) and "The Acceleration Whiplash" (April 2026).
  4. Google DORA, 2024 report and 2025 report.
  5. GitClear, AI Copilot Code Quality research, February 2025.
  6. METR, randomized study of experienced open-source developers (July 2025) and 2026 update.
  7. Gergely Orosz, "Tokenmaxxing as a weird new trend", The Pragmatic Engineer, April 2026.
  8. Reporting on AI-usage mandates: Microsoft (June 2025), Coinbase (August 2025), IT Brew survey (February 2026).
  9. DX, AI Measurement Framework (2025); GitHub, Engineering System Success Playbook (2025).
Kumar Saurabh Johny
Founder and CEO, Darner. Maker of Garth.

J builds Garth, the engineering intelligence platform for AI-assisted development. Before Darner he was CEO of Cathyos, which he built and exited to Automation Anywhere. These are the questions Garth was built to answer, but every metric in this piece can be computed without it.