Three Models, One Point
Google dropped three models yesterday. Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber. Not one pitched itself on intelligence. Flash-Lite runs at 350 tokens per second. Flash 3.6 cuts output token usage by 17 percent. Flash Cyber does security scans and nothing else.
Faster. Cheaper. Narrower. That was the whole pitch, three times in one afternoon.
A year ago every release led with benchmark scores. Longest context window. Hardest reasoning test. The implicit message was always the same: wait for the next one. It’ll be smarter.
Nobody is saying that anymore. The last three major releases from Google, Anthropic, and OpenAI all led with efficiency, not capability. Models got cheaper. They got faster. They got more specialized. The intelligence ceiling stopped being the constraint.
The constraint is the workflow. A 350-token-per-second model is worthless inside a chatbot sidebar. It’s built for teams running dozens of agents across procurement, compliance, and ops — thousands of tasks per hour where cost per interaction is the metric that matters.
If you can ship three models on a Tuesday and compete with yourself, the model isn’t the hard part. The hard part is knowing what to run it inside. That question hasn’t gotten easier. It just got cheaper to answer wrong.