# AI Commit Attribution: How Accurate Is the Cost on a Pull Request?

Any tool can print a dollar figure against a PR. The number can rest on the agent's own git commit call or on a same-day guess, and most dashboards render both identically. How to tell which you have.

_Published 2026-10-02._

---
import TLDR from '@/components/TLDR.astro';
import DefinitionBox from '@/components/DefinitionBox.astro';
import Callout from '@/components/Callout.astro';
import FAQBlock from '@/components/FAQBlock.astro';

<TLDR>
- A dollar figure on a pull request is the output of a join, and the join runs on evidence of wildly different strength. Most dashboards print the strong row and the weak row in the same font.
- The strongest evidence is the agent's own `git commit` tool call matched to the sha it produced. The weakest in common use is a timing heuristic: same person, same day, therefore same AI.
- Confidence is not coverage. Confidence grades one row. Coverage tells you how many rows exist at all. A tool can be highly confident about 11% of your merged PRs.
- Join confidence and cost kind are separate axes. A measured dollar can sit on a guessed join, and a guessed dollar can sit on a certain one.
- A low-confidence row is not a wrong row. The honest move is to label it and keep it, not to drop it or to promote it.
</TLDR>

A pull request page tells you it cost $42.17. Nobody asks the next question, which is where the $42.17 came from, and the answer is different for the PR above it and the one below.

One of those rows might be the measured token spend of a session that ran the `git commit` command which produced the sha. Another might be a person's daily provider total, sliced across the four PRs they happened to touch that Tuesday. Both render as a dollar sign and two decimal places. Only one of them is evidence.

## What can a per-PR cost figure rest on?

Four kinds of evidence, in descending order of what they prove. The useful question for each is not whether it is accurate, which nobody can tell you, but what it can and cannot support.

| What the figure rests on | Label | What it supports | Where it breaks |
|---|---|---|---|
| A `git commit` tool call inside the session, matched to the commit within 30 seconds | deterministic | Per-PR cost, per-session cost, anything you would defend in a review | Commits made outside the agent, from a plain shell |
| A session id written into the commit body by a hook | deterministic | The same, when the id resolves to a session that was actually ingested | A developer who turned the trailer off, and inheritance (below) |
| A git note on the sha naming the session | deterministic | The same, and it survives a rebase that rewrites the message | Nothing writes notes unless somebody opted in |
| An AI co-author trailer, on the session's own branch, inside the session window, with no tool call matching | inferred | A directional read, a trend, an aggregate with its label attached | Two sessions on one branch in one window |

Below all four sits the tier the category runs on, and it has no session in it at all. Take a person's provider spend for a day, spread it over whatever that person merged, call the slice a cost. The arithmetic is defensible in aggregate and it is not evidence about any individual row. A rebase, a cherry-pick, a pairing session or a day of two parallel agents all move the number without anything in the data changing.

<DefinitionBox term="Attribution confidence">
A label stored on the join between an agent session and a commit, not on the commit and not on the dollar. Three values. `deterministic` means the session produced this commit. `inferred` means AI was involved and the timing and branch point at this session. `estimated` means there is no session evidence and the figure is an allocation. Every joined figure carries its label, and a view that aggregates rows reports the mix rather than averaging it away.
</DefinitionBox>

## Why is the weak tier the common one?

Because the two halves of the question are keyed differently and share no column. Spend lands at account, user and day grain. A provider invoice gives tokens per model per day. A seat subscription gives one monthly figure per person and no per-request breakdown at all. Output lands at commit grain: a sha, an author email, a timestamp, a diff. Nothing on the billing side names anything on the git side.

Allocation is what you do when you have both halves and no key. It is a reasonable response to a real problem. The failure is not the method. The failure is rendering its output in the same cell as a measured join and letting the reader assume the two mean the same thing.

Getting off that tier means capturing cost at session grain in the first place, which is a different data source: the transcript the agent already wrote to disk, carrying the tool calls and the per-turn token counts, in the repo it ran in. That is the join key, and it only exists if something was reading the agent rather than the invoice.

## The case that proves the labels are load-bearing

Here is the one I would use to argue for tiers if I had to argue for them once.

A commit trailer is written from the shell environment the commit inherited. Claude Code exports its session id to Bash children, so a hook can stamp `TokenJam-Session: <id>` into the message without guessing. A subagent inherits its parent's session id. So a commit that a subagent's own Bash tool ran carries the **parent's** trailer.

Treat that trailer as first-hand evidence and the parent session absorbs every commit its children produced, at the top confidence, on paper. A fan-out run with eight workers lands all eight workers' output on one session's cost line. The number is wrong in both directions at once: the parent looks productive and expensive, each child looks like it shipped nothing.

The rule that fixes it is small. When a trailer names a session and a tool span in a *different* session matched the same commit, the trailer is inherited evidence and the row is written at `inferred`. The session whose tool call ran the command gets the deterministic row. When nobody else is known to have run the commit, or when the named session is the one that ran it, the trailer stands at `deterministic`.

<Callout type="info">
A tiering scheme that has never demoted anything is decoration. The test of one is whether it has a rule that costs you a strong-looking row, and the inherited-trailer case is that rule here.
</Callout>

## Confidence is not coverage, and the difference is where dashboards go wrong

Confidence grades a row that exists. Coverage counts the rows. A tool can be perfectly confident about 11% of your merged PRs and the two numbers will both look healthy on their own.

Only the top two tiers count toward coverage, which is deliberate. An allocated row is not evidence that a commit was joined, so counting it would let coverage rise by adding guesses. Here coverage is the share of your default-branch commits in the window, on the repos those sessions ran in and authored by the people who ran them, with reverts excluded, that have a `deterministic` or `inferred` row behind them.

Read it before the dollars, because it bounds them. And when there is nothing to cover, the number is a dash rather than a zero. A zero reads as "none of your work was attributable," which is a finding. A dash reads as "this window had no default-branch commit in scope," which is the truth.

## Two axes, not one

![Two-by-two grid: measured and allocated cost down the side, a joined and a guessed attribution across the top, with the measured-and-joined cell the only one fully lit.](/blog/images/ai-commit-attribution-confidence/confidence-cost-axes.jpg)

The third thing to hold separate: how the row was joined and what kind of dollar is sitting on it. A measured cost can ride on an inferred join. An allocated list price can ride on a deterministic one. Collapsing them into a single quality score loses the only information you need to act.

So four combinations, and each needs a different response. Measured cost on a deterministic join is the figure you quote. Measured cost on an inferred join is a figure you trend and do not attach to a person. An allocated figure on a certain join tells you the join is fine and the pricing is a rate card, which is [a separate problem with its own fix](/blog/2026-09-28-ai-cost-estimate-vs-actual-list-price). Allocated on estimated is a model of your spend, and the honest render says so in the cell.

## What do you do with a low-confidence row?

Not delete it. A row at the weaker tier is not a wrong row; its evidence is thinner, and dropping it swaps a stated uncertainty for a silent one. Your totals get smaller, your coverage gets better looking, and the work it represented still happened.

Four moves, in the order they pay off.

**Say which tier it is, next to the number.** One glyph per row and a mix in the aggregate. Everything else on this list is optional.

**Raise the evidence where it is cheap.** A session trailer or a git note turns a window match into a resolvable id. That is a post-commit hook and it is opt-in, because writing to commit messages or to `refs/notes/*` without being asked is not ours to do.

**Keep weak rows out of per-person comparisons.** Inferred evidence is fine for a quarterly trend and unfit for a sentence with somebody's name in it. The co-author trailer the weak tier depends on is a setting a developer can switch off, so a trailer-only pipeline measures the people who left the default alone.

**Ask for the exclusion reason rather than the exclusion.** "Not attributable" is not a reason. "No session was ingested for this author at the time" is, and it tells you what to fix.

Rows that cannot be joined at all are simply absent here. `tj` has no allocation path, so it never writes a row at the `estimated` tier; the label stays defined so that an allocated figure can never turn up unlabelled, and coverage is how you find out how many commits fell outside.

## Checking it on your own history

The join runs locally off the transcripts Claude Code and Codex already write to disk, with read-only git and nothing in your request path.

```bash
pipx install tokenjam
tj onboard              # backfills the transcripts already on disk
tj optimize shipped     # the split, with coverage beside it
```

`tj optimize shipped` prints coverage and the measured cost of each shipped state. The per-row labels show up where the rows do: the commit chips on a session page in Lens carry the confidence glyph, and `GET /api/v1/sessions/{id}` returns `confidence` and `source` on every joined commit. If you want the other half of this, the spend behind sessions that produced no commit at all, that is [unshipped burn](/blog/2026-09-28-agent-session-cost-unshipped).

<FAQBlock items={[
    {
      question: "My tool says a PR cost $42 and my provider bill does not reconcile. Which one is wrong?",
      answer: "Probably neither, and the gap is usually two things stacked. The per-PR figure may be an allocation of a daily total rather than a measurement of that PR's sessions, and the rate behind it may be a public price card rather than what you actually pay. Check the join first, then check the pricing mode."
    },
    {
      question: "Is a Co-Authored-By trailer enough to attribute a commit to an agent?",
      answer: "It is enough to know AI was involved. It is not enough to know which session, because the trailer carries no session id and it is written from an environment a subagent inherits. It also depends on a setting any developer can turn off, so absence of the trailer is not evidence of absence of AI."
    },
    {
      question: "Can I just use the top tier and ignore the rest?",
      answer: "You can, and you should expect coverage to fall when you do. Filtering to deterministic rows gives you a cleaner number over a smaller slice of your history. That is a real trade and worth making deliberately rather than by accident."
    },
    {
      question: "What breaks attribution most often?",
      answer: "Commits made outside the agent. A developer who has the agent write code and then commits by hand from their own shell leaves no tool call and, without a hook, no trailer. The work was AI-assisted and nothing in the data says so."
    }
  ]}
/>

Running coding agents and want the provenance of every cost figure spelled out? `pipx install tokenjam`, then `tj optimize shipped`. Local, no account. The labels are in the output.