Guide
Measuring the impact of AI coding tools.
Measuring AI impact means comparing adoption and usage of tools like GitHub Copilot, Cursor and Claude Code against delivery outcomes — cycle time, pull request size, rework and throughput — for the work those tools were involved in. What that comparison supports is a correlation. It does not, and cannot, establish causation, and the difference matters when the number is going in front of a board.
Start with adoption, because most of the money is there
Before any impact question, answer the boring one: how many of the seats you are paying for are actually being used? Across organisations that measure this for the first time, a meaningful proportion of assigned seats have never been activated, and a further group activated once and stopped. That is a direct, immediate saving and it requires no statistical argument at all.
Adoption also has to come from the provider’s own usage data rather than being inferred from commit patterns. Inferring AI usage from code characteristics is unreliable in both directions, and it produces exactly the kind of number that collapses the first time somebody senior questions it.
Then compare, carefully
The honest comparison is between work that was AI-exposed and work that was not, on measures that are already meaningful: cycle time, pull request size, rework rate, review latency. Report the sample size alongside it, because these comparisons are frequently run on samples far too small to say anything.
The confound is unavoidable and should be stated rather than hidden. Engineers choose when to use AI assistance, and they choose it more readily for certain kinds of work — boilerplate, tests, unfamiliar APIs. So the AI-exposed population is not a random sample of your work, and any difference you observe is partly a difference in the work itself. This does not make the comparison useless. It makes it a correlation, which is what it should be labelled.
A useful second cut is variance rather than average. Some of the strongest observed patterns in this area are not "AI-assisted work is faster on average" but "AI-assisted work is more consistent", or the reverse — and consistency is often the more valuable property.
What to say to the board
The defensible version of the argument has three parts. First, adoption: this many seats, this many active, this much cost, this much waste identified. Second, the correlation: AI-exposed work shows these differences on these measures over this sample, presented as a correlation. Third, what you are doing next: which teams are lagging on adoption, what is being changed, and when you will look again.
The version to avoid is a single percentage productivity uplift. It will be asked to justify itself, the confound will be found within about two questions, and the credibility of every other number you presented goes with it.
Where this measure goes wrong
Each of these produces a number that looks reasonable and is not.
Inferring AI usage from code characteristics
Unreliable in both directions. Use the provider’s own organisation analytics.
Reporting a single productivity uplift figure
It cannot survive scrutiny, and losing it costs you the credibility of everything else in the deck.
Ignoring the selection effect
Engineers choose when to use AI, and for what. The exposed sample is not random and saying so up front is far stronger than being caught.
Blending providers into one adoption number
Copilot, Cursor and Claude expose genuinely different metrics. A combined figure averages incompatible measurements.
Measuring output instead of outcomes
More code, faster, is not the goal. Cycle time to production and rework rate are.
Frequently asked questions
Can you prove AI coding tools improve productivity?
Not from delivery data, and no vendor can. What you can show is whether AI-exposed work differs measurably from other work on cycle time, size and rework, over a real sample, with the selection effect acknowledged. That is a correlation, and it is enough to make an investment case honestly.
What should we measure first?
Seat activation. Assigned seats that were never used are a direct cost with no argument attached, and identifying them usually pays for the measurement exercise several times over.
Which AI tools expose usable organisation data?
GitHub Copilot exposes seats, active users and acceptance. Cursor exposes team usage. Anthropic exposes two different products — a Console organisation reports Claude Code usage and estimated cost, an Enterprise organisation reports seats, active users, sessions, tool actions, lines changed, commits and pull requests but no cost.
Does usage through Bedrock or Vertex show up?
Not in Anthropic’s organisation analytics APIs, which do not include Claude usage hosted through Amazon Bedrock, Google Vertex AI or Microsoft Foundry. If that is where your usage runs, this is a real gap and worth knowing before you build reporting on it.
See measuring ai impact for your own teams
Pacia computes this from your repositories and issue tracker, banded and trended, with a drill-down to the work behind every figure.