Bottom line
Google just cut the price of the workhorse models that run AI agents. If your automations route each job to the best-priced model, your cost of answering, booking, and follow-up keeps falling on its own.
What happened
On July 21, 2026, Google released two lower-cost models in its Gemini line, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, both aimed at running AI agents on high volume without the token costs piling up. Google positions Flash for coding and multimodal work and Flash-Lite for fast, high-throughput jobs like document processing and agentic search, with reported per-token prices well below its flagship tier. The pitch is simple: as businesses move from one-off AI questions to agents that run continuously, the speed and cost of the workhorse model is what decides whether that is affordable.
Reported by AI News. Our analysis is below.
The EMOR AI take
The pattern worth watching is not which lab is ahead this week, it is the direction of the price line, and it points one way: down. Every few weeks another cheaper, faster workhorse model ships, which means the automations we run for clients, lead follow-up, document handling, content, get cheaper to operate over time without anyone changing a thing. We route each job to the model that fits it, so falling prices reach your bottom line instead of a lab’s margin.
There is a reason this matters for a small business specifically. The falling cost of these workhorse models is what turns AI from a one-off novelty into something you can run on real volume, every lead, every document, every day, at a price that still makes sense. The businesses that build that into their operations now compound the advantage as the underlying cost keeps dropping. Put cheap, reliable intelligence to work on your routine volume and let the market keep lowering your bill.
What this means for your business
- Judge AI vendors on whether they pass falling model prices to you
- Move routine, high-volume work onto automation now
- Expect the cost of running AI agents to keep dropping
Frequently asked questions
What are Gemini 3.6 Flash and Flash-Lite built for?
They are Google’s lower-cost models aimed at running AI agents on high volume. Google positions Flash for coding and multimodal work and Flash-Lite for fast, high-throughput jobs like document processing and agentic search, at per-token prices well below its flagship tier.
Why do cheaper AI models matter for a small business?
The model price sets the cost of every automated task you run. As workhorse prices fall, jobs like lead follow-up, document handling, and content get cheaper to operate on real volume, which is what makes always-on automation affordable for a small team.
How does a business actually capture these price drops?
Use systems that route each job to the best model for the task instead of locking everything to one vendor. When any lab cuts prices, a routed system shifts the work and the savings show up in your operating bill without a rebuild.
Read the original report
Google’s Gemini 3.6 Flash targets enterprise agent token costs
AI News
Want this working for your business?
We build the AI receptionists, automations, and content systems businesses actually run on, custom or product.
Or start with your free AI growth roadmap.
Go deeper
All articlesA Billion People Are Talking to AI Instead of Typing. Here Is What That Does to Your Phone.
Google says 63% of Gemini's billion monthly users now ask by voice, and on September 4 it starts replacing Assistant with Gemini on Android. A spoken question returns one recommended business, not ten links. Here is what makes yours the one that gets named, and what to fix before the switch.
SEO & AI SearchChatGPT Sells Ads Now. Should a Local Business Buy Them?
OpenAI's self-serve Ads Manager is open, the spend minimums are gone, and ads run inside the assistant your customers ask for recommendations. Here is an honest read on when paid placement in ChatGPT is worth it for a small business, and the free work that beats it almost every time.