AI Code Rework Cost: When You Pay Twice for the Same Lines
A feature lands on Friday. Someone opens it on the following Wednesday, finds the control flow hard to follow, and rewrites most of it. Two pieces of work, one surviving file.
That has always been true of software, and it has always cost twice. What changed is the rate. An agent can produce several hundred lines in the time it takes to read the ticket, at a cost that rounds to nothing against an hour of salary. Nobody audits that output the way they would audit half a day of a person’s work. The second pass through the same file stops registering as a cost event. It is just Wednesday.
What counts as rework?
The plain version: work you paid for that got undone by later work. In a code repository that has a mechanical shape. A commit adds lines to a file, and within some window other commits delete those lines again.
This is old ground. Relative code churn has been used as a defect-density predictor since at least a 2005 Microsoft Research study, and rework as a share of engineering effort predates version control entirely. The measurement is not new and it is not controversial. The reason to pick it up again is that one of the two inputs got much cheaper and much faster, and a cost that used to be self-limiting no longer is.
Why rework is the quality proxy that needs no model in the loop
Most attempts to put a quality number on AI-generated code end up asking a language model to grade a diff. That has two problems that stack. The grader is the same class of system that wrote the code, and the calibration corpus behind the grade is usually unpublished, so you are asked to trust a score you cannot audit.
Rework asks a narrower question with a checkable answer. Did these lines survive? git knows. The computation is numstat subtraction over a date range, it produces the same answer on every machine, and a developer who disputes the result can read the commits themselves and see exactly which deletions drove it.
How do you compute rework from git?
Here is the full method, because the method is the number. Anything vaguer than this is an opinion with a dollar sign on it.
Start from a commit and its own files. git show --numstat --format= --no-renames <sha> gives the lines each path gained in that commit. Merge commits and binary files contribute zero, which is correct: neither one authored a line.
Open a window from the commit forward. Fourteen days in the implementation I will use for the rest of this post. The window closes at the commit’s timestamp plus fourteen days, or at now, whichever comes first. A commit three days old has not finished being measurable yet.
Count what later commits deleted from exactly those paths. git log --all --numstat --no-renames --after=<commit time> --before=<window end> -- <the commit's paths>, summing the deletion column per path. Scoping to the commit’s own files matters. Deletions elsewhere in the repo are somebody else’s business.
Exclude the author’s own iteration. Commits belonging to the same agent session as the original, and the original commit itself, are skipped. A session that commits, notices a bug, and commits a fix is working. It is not reworking itself, and counting it that way would make every careful session look like a disaster.
Clamp per path. Deletions are capped at the additions. A file can lose more lines than this commit put in there, and the excess belongs to whoever added them.
Then apply a threshold. Sum additions and surviving deletions across all of a session’s commits. If at least half of the lines it added were deleted again inside the window, the session is counted as reworked, and its measured cost joins the rework total.
Four numbers define that measurement: the window, the share threshold, the file scope, and the exclusion rule. Change any one and the figure moves. Which is why the figure ships with a sentence stating all four, how many sessions had line counts to measure at all, and how many of those qualified. A rework dollar with no denominator next to it is unreadable.
What the threshold and the window miss
A heuristic is a set of deliberate blind spots. These are the ones here, and I would rather you hear them from me than find them yourself in a review.
A rename reads as a total loss. With --no-renames, moving a file is a delete of every line at the old path and an add at the new one. A tidy-up that relocates a module will light up as 100% rework of the original commit. The same goes for a lockfile regenerated by a dependency bump, or a formatter sweeping the tree. Lines left, and the method cannot tell you they left in good health.
The threshold is a cliff. Forty-nine percent of a session’s lines deleted counts as nothing. Fifty-one percent counts the session’s entire cost. Two sessions two lines apart land on opposite sides, and the second one contributes its whole dollar figure rather than its share of damaged lines. That is a deliberate choice to keep the number simple, and it makes any single session’s row a weak piece of evidence. The aggregate is what carries signal.
Fourteen days is a guess. Rework on day 20 is invisible to a 14-day window. Published figures in this space often use 30, 60 and 90 day horizons, and a longer window will always find more, because given enough time all code is deleted. There is no principled correct answer. There is only the requirement to say which one you used.
Self-rework hides. The exclusion that stops a session blaming itself also means a long session that wrote a module three times over looks perfectly clean. That spend is real and this metric does not see it. Consecutive edits to one path with identical content are a separate measurement for a reason.
Unjoinable work is not in the denominator. If a commit cannot be tied back to a session, it has no cost to attribute, so it sits outside the measurement entirely. The rework total is a floor over the sessions you could measure, never a total over your repository.
And the method cannot tell you why. It establishes that lines were removed. The reason lives in a pull request conversation, or in somebody’s head. Any tool that reports a reason here inferred it, and you should ask how.
Rework is not failure, and a number that implies otherwise is broken
This is the part worth getting right, because the metric is easy to weaponise and useless once you have.
Requirements change after code is written. That is the normal condition of building software for people who are still learning what they want. Review catches real problems, and a reviewer who sends back half a diff is doing the job. A spike exists to be thrown away; its entire purpose is to produce a deletion. An experiment that answered its question and got removed was a cheap success.
So a high rework share is a question, not a verdict. Three readings fit the same number, and they want different responses:
- Work is being started before it is specified. The fix is upstream of the code.
- The agent is producing plausible structure that does not survive contact with the codebase. The fix is in the context it was given, or in the model.
- The team is refactoring aggressively on purpose. No fix. Note it and move on.
Which one you are in is not in the data. You get the figure, you get the most-reworked paths, and then somebody who knows the codebase reads them. A metric that comes with its own conclusion attached is selling you something.
Measured spend, not a saving
One more distinction, and it is the one most likely to get blurred by whatever dashboard you end up looking at.
A rework dollar is money that already left. It is the measured cost of sessions whose output did not survive, priced from the tokens those sessions actually consumed. Nobody is going to give it back. There is no setting to flip that recovers it, and no honest way to put it in a projected-savings column.
That makes it a different species from an optimization figure. “Switch this workload to a cheaper model and the next thousand runs cost less” is a forward claim with a lever attached. “These sessions cost $340 and their lines are gone” is a historical fact with no lever, and its value is diagnostic. It tells you where to look. The reason to keep the two apart on the page is that mixing them produces a savings total nobody can act on, which is how cost tooling loses its reader’s trust.
Which is also why the rework line sits outside the recoverable-waste rollup in tj, carries no projected-saving field at all, and prints its window, its threshold and its denominator every time it appears.
Looking at it on your own history
If you run coding agents and the repository is on your machine, this is computable today without instrumenting anything new. The agent already wrote its transcripts to disk, and git already has the line counts.
pipx install tokenjam
tj onboard # reads the transcripts already on disk
tj optimize shipped # rework line, basis string, most-reworked paths
The rework line arrives with the sentence that states the 50% share, the 14-day window, the numstat method and the session count behind it, plus the top reworked paths by measured cost. Read the paths before the dollar. A single file taking most of the rework is a design conversation; rework spread thin across the tree is usually a specification problem. The other half of the same finding, spend on sessions that produced no commit at all, is unshipped burn, and it needs the same caveat for the same reason.
- Is high code churn after AI-generated commits actually a problem?
- It depends entirely on what caused it, and the churn number cannot tell you. Refactoring and a reviewer rejecting a bad approach look identical in numstat output. Treat a rework share as a prompt to go read the most-affected files rather than as a defect count. The one reading that is solid is the cost: that spend happened and those lines are gone.
- What window should I use for measuring rework?
- There is no correct answer, only a declared one. Shorter windows (14 days) catch work undone while it is still fresh, which is the signal closest to a specification or context problem. Longer windows (30, 60, 90 days) catch more, and much of what they add is ordinary maintenance, because over a long enough horizon every line gets deleted. Pick one, print it next to the figure, and do not compare a 14-day number to somebody else's 90-day number.
- Does a rename or a reformat count as rework?
- With a plain numstat method, yes, and that is a known false positive. Git sees a move as a delete at the old path plus an add at the new one, and a formatter rewriting a file deletes every line it touches. Rename detection reduces it and does not eliminate it. If a path at the top of your most-reworked list was recently moved or reformatted, that is your explanation.
- Can I tell whether the AI or a human wrote the lines that got deleted?
- You can tell which agent session produced a commit if something was watching the agent when it ran, through the agent's own git commit tool call or a session trailer in the message. You cannot reliably get there from commit metadata alone, because a co-author trailer is a setting anyone can turn off and absence of it is not absence of AI. The confidence of that join is a separate question worth asking before you trust a per-session figure.
- Is rework cost recoverable? Can I budget against it?
- No. It is measured historical spend on work that did not survive, with no switch to flip that returns it. Keep it out of any savings total. Its use is diagnostic: it points at the files and the kinds of task where your agent output is not sticking, so you can change something upstream. The money itself is gone.
Running coding agents and want to see which of your spend produced lines that are no longer there? pipx install tokenjam, then tj optimize shipped. Runs locally, no account, and every figure prints the window and threshold it was computed with.