Agent Session Cost: What the Sessions That Left No Commit Cost You
Somebody upstairs has asked what the coding agents are returning, and the honest answer is that you do not have the denominator yet. You can get a token total out of the provider console. You can count merged PRs out of git. Divide one by the other and you have a number that survives a slide, right up until the first person asks the obvious follow-up: how much of that went into work that never landed?
That question has no answer in most stacks, and the reason is structural.
Why can’t your billing data answer this?
Because the two halves are keyed differently and share no join column.
Spend arrives at account, user and day grain. A provider invoice gives you tokens per model per day. A seat subscription gives you a flat monthly figure per person and no per-request breakdown at all. Output arrives at commit grain: a sha, an author, a timestamp, a diff. Nothing in the first set names anything in the second.
The standard workaround is apportionment. Take a person’s daily spend, spread it across whatever that person touched that day, and roll the slices up per pull request. It produces a cost for every PR, and the arithmetic is fine in aggregate. What it cannot do is tell you what share of the total bought nothing, because a session that produced no commit has no PR to be apportioned to. Its cost does not disappear. It gets quietly folded into the PRs that did land, which read a little more expensive than they were, and the discarded work reads as free.

An unshipped session is not a wasted session
This belongs above the figure. A caveat in small print under a dollar sign has never stopped anyone reading it as waste.
Plenty of agent work is supposed to end without a diff. You point an agent at an unfamiliar service and ask it to explain the request path. You have it review a branch somebody else wrote. You use it to chase a production incident at 2am and the fix turns out to be a config change in a console. You ask it to talk you out of an approach, and it does. All of that is work, some of it is the most valuable work of the week, and none of it commits anything.
So an unshipped figure is a measurement, not a verdict. The useful reading is directional: whether the share is drifting up across quarters, and whether one repo or one kind of task keeps ending in nothing. A team whose unshipped share is 30% might be doing a lot of investigation. The same team at 30% with every unshipped session being a retried refactor that kept failing the same test is a different story, and the only way to tell the two apart is to open the sessions and read them.
What does a session-to-commit join actually look like?
The join has to run on evidence, and the evidence is not all the same strength. Four sources, in descending order of how much they prove:
| Evidence | Confidence | What it is |
|---|---|---|
The agent’s own git commit tool call, within 30 seconds of the commit | deterministic | The session literally ran the command that made the sha |
| A session trailer in the commit body | deterministic | The commit names a session that was ingested |
A git note under refs/notes/ai or a sibling ref | deterministic | A note on the sha names the session |
| An AI co-author trailer on the session’s own branch, inside the session window | inferred | Timing and branch agree, nothing stronger exists |
The weakest tier in common use across this category is weaker than any of those: same person, same day, therefore same AI. That is a heuristic wearing a fact’s clothes, and it degrades badly on the day somebody rebases or cherry-picks.
Two details matter more than they look. First, a trailer is a line of text in a commit message, and any developer can turn the setting off, which means a trailer-only pipeline measures the people who left the default on. Second, a trailer that names a different session than a tool call does is inherited evidence. It travels. A subagent carries its parent’s session id, so a commit the subagent itself ran still carries the parent’s trailer. Recording that pair at the top confidence would let a parent absorb all of its children’s output. It gets recorded as inferred instead.
Coverage is the number to read first
Coverage is the share of your default-branch commits in the window that are joined to a session at all, scoped to the repos those sessions ran in and the people who ran them, with reverts excluded.
Read it before you read the dollars, because it is the bound on them. At 20% coverage, a shipped-versus-unshipped split is a sample presented with the confidence of a total. It does not mean the other 80% was unproductive. It means the tooling could not see it. A figure published without its coverage is the single easiest way for an ROI dashboard to be confidently wrong, and it is worth asking any vendor for the number before you look at anything else they show you.
Coverage also moves. Commits made from a plain shell carry no session evidence until something puts it there, which is what tj init --hooks is for.
How to measure it on your own sessions
TokenJam does the join locally, off the transcripts Claude Code and Codex already write to disk, with read-only git and no proxy in your request path.
pipx install tokenjam
tj onboard # backfills the transcripts already on disk
tj optimize shipped
The finding reports how many of the window’s sessions shipped a commit to the default branch, the measured cost of each state, coverage, the largest unshipped sessions so you can open them, and rework cost for sessions whose commits had at least half their added lines deleted again within fourteen days by other work. Sessions with no repo context are counted and disclosed on their own line, because they were never analysed.
It prints this every time, on every surface, because the sentence is doing real work:
Unshipped is measured cost of sessions with no joined commit; a session can ship value without a commit (research, review, ops). Review before acting.
The finding carries no recoverable-dollars field and sits outside the savings rollup on purpose. It belongs to the measurement half of the product, and nothing in it is a proposal. If you want the sessions that genuinely have money in them, where your AI agent bill goes covers the waste patterns that do come with a fix attached.
- What counts as an unshipped agent session?
- Per session, it is a state of its own: no commit joined by any evidence source, distinct from a session whose commits exist but have not reached the default branch, and from one whose commits were all reverted. The rollup is coarser than the states are. It splits spend three ways, shipped, committed and unshipped, and a reverted session lands in that last bucket, because the money bought something nobody kept. So read the unshipped figure as spend that produced nothing you still have, and the per-session state when you want to know which of the two reasons applies.
- Is unshipped spend money I can get back?
- No, and treating it that way will burn your credibility with your own engineers. It is measured cost on sessions that left no commit, and a large share of agent work is supposed to leave no commit: reading unfamiliar code, reviewing someone else's branch, debugging an incident, deciding not to build something. Use it as a trend, and as a pointer to specific sessions worth reading. It is not a savings target.
- Why can't I get this from my provider's billing export?
- Because the export is keyed by account, user and day, and a commit is keyed by sha, author and timestamp. There is no column the two share. You can apportion a daily spend total across the work a person touched, which gives every PR a cost, but apportionment has nowhere to put a session that produced no PR, so that spend gets folded into the ones that did. Getting a real figure needs cost recorded at session grain in the first place.
- How accurate is the session-to-commit attribution?
- It varies by evidence, which is why every row carries a label. The strongest case is the agent's own git commit tool call inside the session, within thirty seconds of the commit. Next are a session trailer in the commit body and a git note naming the session. Weakest is an AI co-author trailer on the right branch inside the session window, recorded as inferred. Read the coverage figure alongside any of it: it tells you what share of your merged work was joinable at all.
- What is a normal unshipped share?
- There is no published benchmark worth quoting, and anyone offering you one is guessing across teams whose work looks nothing like yours. A research-heavy month and a feature-delivery month will differ by a lot on the same team. The number that means something is your own trend over several windows, and the specific sessions behind any jump.
- Does measuring this mean tracking individual developers?
- It does not have to, and it should not by default. The unit here is a session and a repo. A per-engineer ranking built on a commit trailer that any developer can switch off mostly measures who left the setting alone. A measurement system people have a reason to game is worse than no measurement system.
TokenJam is a local-first, OTel-native cost layer for AI agents. No cloud, no signup. pipx install tokenjam, then tj optimize shipped to see what your own sessions left behind and what the rest of them cost.