#token-cost
4 posts
-
Missing Token Counts on Streamed LLM Calls (and Why Your Spend Total Reads Low)
A streamed response reports its token usage in one final payload. When that payload never arrives, the call is recorded with no tokens and prices at zero, so it drops out of every spend total while still counting as a call.
-
Why the Cheapest LLM Can Cost You the Most
A lower cost per token does not mean a lower bill. Here is how a cheaper model runs up more spend through retries, verbose output, and reasoning tokens, and how to compare LLM cost by the finished task instead of the token.
-
Why AI Bills Keep Rising While Token Prices Fall
Token prices drop about 10x a year, yet AI spending is forecast to hit $2.59 trillion in 2026. This is the Jevons paradox for AI, and the fix is measuring your own consumption.
-
AI Pricing Models: Why Token Cost Breaks Usage-Based and Subscription SaaS
When the unit of work is a metered token with a variable, often invisible cost, flat subscriptions and usage tiers stop mapping to value. The case for outcome-based pricing, and why it needs token-cost measurement first.