AI Cost Estimate vs Actual: Is That a Bill, or a List Price?
Someone asks what your coding agents cost last month. You open the dashboard and read off $4,180. Then finance opens the actual statement. Twelve seats at a flat monthly rate, one modest API line for a batch job, and no row anywhere that resembles $4,180. Neither number is a mistake. They are answers to different questions, and only one of them is a bill.
What is the number in your AI cost dashboard actually measuring?
Three quantities hide behind the same dollar sign.
Metered API spend. You send tokens to an API key, the provider counts them, multiplies by a contracted rate, and sends a statement. This is the honest case. The dashboard number and the invoice line describe the same event, and they can be reconciled to the cent.
An amortised seat fee. You pay a flat monthly price per developer. That is real money and it appears on a statement. It does not vary with tokens. Spreading it across usage to get a per-session figure is an allocation, which is a defensible thing to do and a different thing from measuring.
A list-price valuation of subscription traffic. Your developers work under a flat-fee plan. A tool counts their tokens anyway, multiplies by the provider’s public API rate card, and prints the product as spend. Nothing here was billed. The tokens were covered by the fee you already paid.
The third one is the problem, because it produces the biggest number and looks exactly like the first.
Why are published rates the wrong denominator for a subscription?
Because a seat is not a discount on tokens. It is a different product.
Under metered API billing, every token has a marginal price. The thousandth call of the day costs the same as the first, and a 20% reduction in tokens is a 20% reduction in the line item. Under a subscription, the marginal cost of the next token is zero until you hit the cap, and at the cap it is not money at all. It is a rate limit. You wait, or you queue the work for the next window.
So the constraint changes shape. On API billing the scarce resource is dollars. On a seat the scarce resource is throughput, and the honest unit is share of quota. Pricing a seat’s tokens at the rate card converts a throughput problem into a currency nobody in the transaction is using.
Watch the arithmetic pull apart. These figures are illustrative; substitute your own plan and your own volumes.
| Twelve developers on flat-fee seats | |
|---|---|
| Plan cost on the statement | $2,400 / month |
| Tokens consumed | 620M |
| Those tokens valued at a public rate card | $3,900 |
| Amount actually charged per token | $0 |
| The number a rate-card dashboard reports | $3,900 |
The $3,900 is not fabricated. Run the same traffic through an API key and that is roughly what you would pay. As a statement about what happened to your bank account last month, it is off by $3,900.
Where does the list price get typed in?
Look for the settings page. A cost surface that cannot see your contract has to get the rate from somewhere, and there are only two places it can come from: a value you enter, or a rate card the vendor ships and refreshes.
Neither is your agreement with your provider. An entered rate is a guess that ages the moment prices move or your commitment changes. A shipped rate card is a public list price, which is the one number guaranteed not to be what a company with a negotiated contract pays.
There is a second path that is easier to miss. Some coding agents emit a cost figure in their own telemetry. Claude Code’s OTel api_request event carries a cost_usd attribute on every request, which we read in tokenjam/otel/semconv.py. That figure is computed by the agent from token counts at published rates, and the agent does not know whether the account behind it is metered or covered by a plan. Pipe it into a dashboard, sum it, and you have a precise-looking monthly spend total for an account that may have been charged a flat fee the whole time.
How do you check this on your own bill?
Four steps, and none of them need a tool.
- Open the actual statement for the same month the dashboard is reporting. Not the usage console. The invoice.
- Find the per-token line. If your usage sits under a seat or a plan, there will not be one. That absence is the answer.
- Compare the totals. A dashboard figure that exceeds a flat-fee invoice by a multiple is not an overage. It is a valuation.
- Ask the tool what rate it used, and whether it can tell the two kinds of traffic apart. The second half matters more than the first. A tool that knows your plan can label the figure. A tool that only knows a rate cannot, no matter how accurate the rate is.
Most teams run both kinds of traffic at once: seats for interactive coding work, API keys for batch jobs and CI. A total that mixes a valuation and a charge in the same sum is not wrong in one place. It is uninterpretable throughout.
What does an honest render look like?
The fix is not to hide the dollar. A dollar figure is the unit people reason in, and a token count does not translate into a decision. The fix is that the dollar never travels alone.
Here is how the open-source tool handles it, and every part of this is in the source.
Plan tier rides on the session. tokenjam.plan_tier is a column on SessionRecord rather than a span attribute, because a plan does not change call to call inside a session. The ingest pipeline reads the declared plan for the session’s billing account and stamps it there. Valid values run api, pro, max_5x, max_20x, plus, team, enterprise, local, unknown.
Pricing mode is derived from it, not stored. SessionRecord.pricing_mode maps the tier to one of four rendering modes, first match wins: local, then subscription for anything in SUBSCRIPTION_PLAN_TIERS, then api, then unknown. That derivation lives in one place, so no renderer gets to reinvent it.

Every dollar carries a framing block. core/framing.py computes a framing record for a window and the local REST API emits it into responses, so anything reading the API gets the rules alongside the numbers instead of re-deriving them. On a subscription-dominant window that record carries: “Dollar figures price all traffic at API list rates, so they are a list-price equivalent, not an amount billed.” When the plan tier is unknown for every session it says the figures may overstate actual cost, with a pointer at tj onboard --reconfigure. On local inference there is no dollar at all, because there is no marginal cost to price.
Worth being exact about where that text goes, because it is a wire contract rather than a banner. The CLI deliberately does not print it: four commands carry the same comment recording the product decision, no differentiated messaging between subscription and API users. The dashboard surfaces it as a per-figure tooltip reading “Estimated from token usage at API list rates”, a plan-tier badge, and a Spend tile that becomes “Implied plan value” on a subscription rather than a spend number. The qualifier is what a consumer of the API is handed so it cannot render a bare dollar by accident.
Where a dollar would mislead outright, the unit changes. tj quota-audit reports the share of your premium-tier quota that went to sessions a cheaper model could have shaped, as a percentage of premium tokens. Its own module docstring says why it refuses a dollar saving: the subscription majority is on a flat fee, so dollar framing mis-targets them. tj context and tj tokenmaxx do the same at a smaller grain, rendering “X% of cycle tokens” when the pricing mode is subscription.
None of this makes the number more precise. It makes it legible. You can act on “this would cost $3,900 at retail and you were charged a flat $2,400” in a way you cannot act on “$3,900.”
Common questions
- My dashboard says my team spent $4,000 on Claude Code but the invoice is a flat fee. Which is right?
- Both, for different questions. The flat fee is what you were charged. The $4,000 is almost certainly your token usage valued at the public API rate card, which is what the same work would have cost on metered billing. Check whether the tool has any concept of your plan tier; if its settings page has a field for the token rate, the figure is a valuation rather than a bill.
- Why can't a cost tool just read my real rate?
- Because the rate is in a contract, not an API. Providers expose usage, and a negotiated rate or a committed-spend discount lives in a commercial agreement the tool cannot see. So tooling falls back to a published rate card or a number you type in. That is a reasonable fallback. Presenting its output as an invoice amount is where it goes wrong.
- Is the list-price number useless then?
- No, it is the right number for two real questions. What would this workload cost if we moved it to metered API billing, and how much is this seat actually worth to us. That comparison is worth having. Calling it spend is where it goes wrong. It is the wrong number for reconciling against finance or forecasting next quarter's bill, which is what people usually reach for it to do.
- What should I measure instead if my team is on subscription seats?
- Share of quota. On a flat-fee plan the scarce resource is throughput inside the window, not dollars, so the question that changes a decision is which sessions and which models are eating the cycle. Run pipx install tokenjam, then tj quota-audit to see what share of your premium-tier quota went to sessions a cheaper model could have handled. It reports in premium token share rather than dollars, deliberately.
- How does TokenJam decide which framing to use?
- Each session is stamped with the plan tier declared for its provider during tj onboard, and a pricing mode is derived from that tier: local, subscription, api, or unknown. Dollar figures still render for subscription accounts, because dollars are the unit people reason in, and the framing record the API emits alongside them states they are a list-price equivalent rather than an amount billed. On screen that arrives as a per-figure tooltip and a plan badge rather than a banner, and the CLI prints no billing-mode line at all by an explicit product decision. When no session has a known plan tier, the record says the figures may overstate actual cost.
Further reading
- Quota, not cost: why /cost is the wrong number on Claude Max. The same problem at one provider’s scope, with the plan window as the unit that replaces the dollar.
- Why AI bills rise as token prices fall. What happens to the rate-card denominator when the rate card keeps dropping.
TokenJam is a local-first, OTel-native cost layer for AI agents. No cloud, no signup. pipx install tokenjam, then tj cost to see your figures with the plan-tier framing attached.