ACTIVE  ·  BUILDING  ·  v1.0 2026-07-21  ·  JL:IOTA:001
No. 091 · 2026-07-12

The Bill Nobody Modeled

DISPATCH  ·  LOGGED WITH MAI

Companies planned AI budgets in 2024. Most assumed inference costs would fall 80% by now. They were right about the price per token. Wrong about everything else.

Usage didn’t stay flat. Once an agent works, people route more through it. A customer support agent that handled 200 tickets a day in January handles 1,400 by July — not because it got better, but because three departments discovered it existed and started piping work to it. The per-unit cost dropped. The total bill tripled.

Then there’s the infrastructure nobody accounted for. Observability. Evals. Guardrails. Retrieval pipelines that need their own compute. One mid-market SaaS company I spoke with spends more on their RAG infrastructure than on model calls. The model is the cheap part. Everything around it is where the money goes.

The real problem: AI costs don’t behave like SaaS seats. They scale with activity, not headcount. Finance teams built forecasts on per-seat logic. That model breaks when one agent can generate thousands of API calls in an hour based on unpredictable demand.

The companies getting this right treat AI spend like cloud compute — metered, monitored, with circuit breakers. Not like a software license you budget once and forget.

LOGGED WITH MAI  ·  2026-07-12  ·  No. 091
← All Dispatches