Anthropic released Claude Sonnet 5.5 on September 28, saying it produces output more than 30% faster than Sonnet 5 and costs up to 30% less per task in its own testing. The price per unit of use did not fall. The model generally does the work in fewer steps. That distinction is the business story: if you want to know whether AI is getting cheaper, measure the cost of a finished, checked task, not the price printed on a model page.
A lower bill for generating text is not a win if someone spends the savings repairing the result.
Pick one recurring task. Decide what finished means before measuring speed. Otherwise the mistakes simply land on the person checking the work, and the savings look better on paper than they feel at the desk.
Where the apparent savings go
Take a support reply. The software writes an answer quickly. An employee checks the customer’s account, corrects a mistaken policy reference, rewrites the tone, and sends it. The model’s bill might be tiny. The employee’s time is not. A draft that arrives twice as fast but still needs that repair has not saved much work.
But if the draft gets the account detail and the decision right, review becomes a check rather than a reconstruction. That saves time even if the vendor charges the same price per unit of use.
Anthropic’s launch includes an early-testing account from Zendesk: across hundreds of real support use cases, Zendesk says Sonnet 5.5 made fewer wrong decisions and processed tickets 20% faster than the Claude models it uses in production. That is a vendor-published customer statement, not a guarantee that another support desk will see the same result. It does show the right unit of analysis. Zendesk looked at the ticket, not just the model response.
The same accounting applies to a finance exception or a vendor comparison checked against contract terms. Who had to pull the source documents back up? Count that time.
The price did not tell the whole story
Anthropic lists Sonnet 5.5 at the same published prices as Sonnet 5: $2 per million units of incoming text and $10 per million units of generated text. It says the new version typically needs less text and fewer steps to complete the same job, which is why its measured cost per task can be lower. These are Anthropic’s figures, drawn from its release testing. The improvement will vary with the work.
The company’s examples also make the limits clear. It positions Sonnet 5.5 for well-scoped everyday work and Opus 5.5 for complex, open-ended work that needs sustained judgment. That is a more useful distinction for a team lead than a leaderboard. You do not need your strongest possible assistant on every routine draft. You also should not assign an ambiguous, high-stakes decision to the cheaper option solely because its first answer looks polished.
Match the job to the judgment it requires. AI can prepare a document against an agreed template. A person should decide whether to approve a disputed customer exception. If the policy is unambiguous and the exception routine, test a wider scope later. Keep the handoff visible.
That is why simplicity matters in AI implementation. A narrow task with a clear finish line is easier to price, easier to review, and easier to improve than a vague mandate to “use AI more.”
Run the comparison on your own work
Choose a recurring task with enough volume to observe, such as a weekly operating report or a set of incoming support tickets. Write down what a good result must contain and which mistakes make it unusable. Then take a sample of past tasks and run the current method and the proposed AI-assisted method through the same finish line. Do not remove the human check for the test.
Record four things: time to a usable result, employee review and repair time, rework after delivery, and direct software cost. If the new method changes who performs the work, include the time at the handoff. A polished draft that sends an analyst back to source documents is not finished. Neither is an answer that a customer later has to correct.
A shared sheet with one row per completed task will expose the pattern. Separate tasks by difficulty. Keep examples of failures; they may explain more than the average completion time.
Anthropic quotes Slack as saying Sonnet 5.5 improved nearly all of its offline Slackbot evaluations while using about 14% less generated text than Sonnet 5. That is promising, but it is still an offline test. The question for a company running Slackbot is what happens when a real person asks for help during a busy workday. Did they get the answer they needed? Did they have to ask again? Did anyone have to clean up behind it?
The release is a good reason to rerun a bounded trial. Use work you understand and the same acceptance standard you would apply to a colleague. If the new method is faster and cheaper after review, expand the scope. If it merely produces more drafts, the bottleneck is somewhere else.
I would rather see a report of completed work than a chart of tokens purchased: what passed review, what it cost, and where a person still had to step in.
Research and structure: Mai. Editorial direction and voice: John Lipe.