Open the vendor dashboard for any AI coding tool and you’ll see acceptance rate near the top. It’s the number managers quote in staff meetings. GitHub’s own research puts it at roughly 30% of suggestions accepted, and the Copilot metrics API gives you suggestion and acceptance counts per editor, language, and day. So a team sees 30%, the line goes up and to the right, and the renewal gets signed.

That number measures a keystroke. Finance is paying for software that stays shipped. The output measure that belongs on the dashboard is the one no vendor shows by default: of the lines your team accepted, how many were reverted or rewritten within 30 days?

Acceptance rate was validated against how developers felt

Acceptance rate didn’t become the headline metric by accident. The GitHub study in Communications of the ACM tested several usage signals and found acceptance rate was the best predictor of perceived productivity. Perceived is the important word there. Developers filled out surveys, and the survey scores tracked how often they hit Tab.

That’s a real finding. It’s also a finding about sentiment. It never claims the accepted code survived review, passed the next sprint, or stayed out of the incident channel. And GitHub sells the tool. Treat the 30% as a vendor number about engagement, not about output.

The newer tooling has the same blind spot one layer up. Dash0’s AI Coding Insights lets you sort every Claude Code session by cost or error count and read the full transcript. That’s useful for debugging an agent. But cost and errors-per-session both describe what went into the work. Neither tells you what happened to the code after the session closed.

Rework is where AI-assisted code shows its real cost

The independent evidence points the same direction. GitClear analyzed 153 million changed lines and projected code churn, meaning lines reverted or updated within two weeks of being written, to double in 2024 against its 2021 pre-AI baseline. GitClear sells code analytics, so discount it too. But the 2024 DORA report, which has no product to sell here, found that a 25% increase in AI adoption was associated with an estimated 7.2% drop in delivery stability. And METR’s randomized trial found experienced open-source developers were 19% slower with AI tools while believing they were 20% faster.

Put those next to the CACM result and you get the core problem. The metric vendors chose tracks how developers feel. The independent studies that measured outcomes found instability, rework, and a perception gap. Acceptance rate sits on the wrong side of that gap.

The $80K AI coding pilot post priced the senior review hours on the cost side. This is the matching hole on the output side. A suggestion that gets accepted on Tuesday and rewritten on the following Monday shows up on the dashboard as a win. Then it shows up again as new work.

You can measure 30-day rework from git history you already have

You don’t need a vendor for this. Git already records who wrote every line and when it changed. The method:

  1. Take every non-merge commit that is at least 30 days old.
  2. Tag it AI-assisted or not. Claude Code adds a Co-Authored-By: Claude trailer to its commits by default, so the tag comes free.
  3. Count the lines each commit added.
  4. Run git blame at the repo state 30 days later and count how many of those lines are still attributed to the original commit.
  5. Rework rate = 1 − (kept ÷ added), computed separately for each cohort.
#!/usr/bin/env python3
"""30-day rework: share of added lines gone or rewritten 30 days later."""
import subprocess
from collections import Counter

DAY = 86400

def git(*args):
    return subprocess.run(["git", *args], capture_output=True,
                          text=True, check=True).stdout

def added_lines(sha):
    files = {}
    for row in git("show", "--numstat", "--format=", sha).splitlines():
        added, _, path = row.split("\t", 2)
        if added != "-" and int(added) > 0:   # skip binaries
            files[path] = int(added)
    return files

def kept_lines(sha, path, rev):
    try:
        out = git("blame", "-w", "-M", "-C", "--line-porcelain", rev, "--", path)
    except subprocess.CalledProcessError:
        return 0   # file deleted or moved out of reach: counts as reworked
    return sum(1 for line in out.splitlines() if line.startswith(sha))

def rev_after(sha, days):
    ts = int(git("show", "-s", "--format=%ct", sha))
    return git("rev-list", "-1", f"--before={ts + days * DAY}", "HEAD").strip()

def main(since="120 days ago"):
    log = git("log", "--no-merges", f"--since={since}",
              "--until=30 days ago", "--format=%H%x09%B%x00")
    totals = {"ai": Counter(), "human": Counter()}
    for entry in filter(None, (e.strip() for e in log.split("\x00"))):
        sha, body = entry.split("\t", 1)
        cohort = "ai" if "co-authored-by: claude" in body.lower() else "human"
        rev = rev_after(sha, 30)
        for path, n in added_lines(sha).items():
            totals[cohort]["added"] += n
            totals[cohort]["kept"] += kept_lines(sha, path, rev)
    for cohort, c in totals.items():
        if c["added"]:
            print(f"{cohort}: {c['added']} added, "
                  f"{1 - c['kept'] / c['added']:.1%} reworked by day 30")

if __name__ == "__main__":
    main()

The -w -M -C flags tell blame to ignore whitespace and follow lines that moved within or between files. Without them, a refactor that only reformats or relocates code gets counted as rework. The 30-day cutoff on --until matters too. A commit from last week hasn’t had its 30 days yet, and counting it would make every cohort look stable.

The objection: some rework is just healthy iteration

Here’s the best case against this metric. Code gets rewritten because teams learn. A spike gets thrown away on purpose. A first pass at an API gets reshaped once the second caller shows up. If you punish 30-day rework, you punish exploration, and engineers will learn to stop committing early.

That objection has a specific mechanism, and the two-cohort design deals with it. Healthy iteration happens to human-written code too, in the same repo, under the same review culture, against the same product churn. Reading a single rework rate on its own tells you nothing. The signal is the difference between the AI cohort and the human cohort over the same window. If both sit at 15%, AI isn’t adding rework and the acceptance-rate story holds. If the AI cohort sits at 30% against a human 15%, AI-assisted code is being rewritten at twice the human rate, and the vendor dashboard counted every one of those lines once, as a win.

Don’t turn it into a per-engineer scorecard, either. It’s a cohort measure for a renewal decision. Once it’s used to grade people, people start gaming the trailers.

Where this measurement breaks

The script has real limits, and you should know them before you put the output in front of a CFO.

Squash merges erase the attribution. If your team squash-merges PRs, the trailer survives only when it lands in the squash message. Many merge tools drop it. Check a few merged PRs by hand before trusting the cohort split.

Inline completions leave no trailer. Copilot-style Tab completions get mixed into human commits without any marker. For those tools, compare whole-repo rework over the 90 days before rollout against the 90 days after. It’s noisier, and it’s still more honest than acceptance rate.

Generated files distort the count. Lockfiles, snapshots, and codegen output get rewritten by machines on every run. Exclude them with a pathspec, or the cohort with more dependency bumps will look worse for reasons that have nothing to do with AI.

Thirty days is a choice. GitClear uses two weeks. Thirty days catches the rewrite that comes when the second caller shows up, which two weeks often misses. Pick a window, write it down, and don’t change it between quarters.

None of these limits rescues acceptance rate. Every one of them is a reason to measure carefully. None is a reason to go back to counting keystrokes.

The number to bring to the renewal meeting

Go back to the dashboard. It says 30% of suggestions were accepted. That figure was validated against how developers felt, and the vendor selling the seats reports it.

Ask the question finance should be asking: of that 30%, what fraction was still in main on day 30, and how does that compare to the code your engineers wrote without help? It comes down to two cohorts and one ratio, kept ÷ added, and you can compute it from your own git log this afternoon.