The Bill Nobody Modeled
Companies planned AI budgets in 2024. Most assumed inference costs would fall 80% by now. They were right about the price per token. Wrong about everything else.
Usage didn’t stay flat. Once an agent works, people route more through it. A customer support agent that handled 200 tickets a day in January handles 1,400 by July — not because it got better, but because three departments discovered it existed and started piping work to it. The per-unit cost dropped. The total bill tripled.
Then there’s the infrastructure nobody accounted for. Observability. Evals. Guardrails. Retrieval pipelines that need their own compute. One mid-market SaaS company I spoke with spends more on their RAG infrastructure than on model calls. The model is the cheap part. Everything around it is where the money goes.
The real problem: AI costs don’t behave like SaaS seats. They scale with activity, not headcount. Finance teams built forecasts on per-seat logic. That model breaks when one agent can generate thousands of API calls in an hour based on unpredictable demand.