#llm-cost
3 posts
-
Missing Token Counts on Streamed LLM Calls (and Why Your Spend Total Reads Low)
A streamed response reports its token usage in one final payload. When that payload never arrives, the call is recorded with no tokens and prices at zero, so it drops out of every spend total while still counting as a call.
-
AI Cost Estimate vs Actual: Is That a Bill, or a List Price?
Most AI cost dashboards price your tokens at a published rate card, including for people on flat-fee seats. Here is why that number cannot be reconciled against an invoice, and what to check on your own bill.
-
Why the Cheapest LLM Can Cost You the Most
A lower cost per token does not mean a lower bill. Here is how a cheaper model runs up more spend through retries, verbose output, and reasoning tokens, and how to compare LLM cost by the finished task instead of the token.