If you've been holding off on AI agents because the token bills scared you, Google just handed you permission to stop waiting. The company's released three new Gemini models designed to run autonomous AI agents at a fraction of what you're paying now — with one model cutting costs by up to 65% on complex engineering tasks. And that's before their flagship Gemini 3.5 Pro arrives later this year.

For Australian business owners, this isn't academic. It means the AI automation you've been pricing at $500–$2,000 per month might suddenly cost $200–$700. It means you can actually afford to let an AI agent handle your customer data processing, your invoicing workflows, or your technical troubleshooting — 24/7, without flinching at the bill.

What happened?

Key Takeaways
  1. Google's Gemini 3.6 Flash cuts token costs by up to 65% on long-horizon tasks — making AI agents dramatically cheaper to run at scale, particularly for complex engineering and multi-step workflows Source: Google DeepMind
  2. Three models released simultaneously — Gemini 3.6 Flash (the workhorse), Gemini 3.5 Flash-Lite (stripped down for speed), and Gemini 3.5 Flash Cyber (security-hardened) give you options for different workloads and budgets
  3. Flagship Gemini 3.5 Pro is coming this year — expected to deliver even more capability without proportional cost increases, making this a good moment to lock in pricing with Google before rates shift
  4. This solves the AI agent affordability problem — for SMEs that have watched AI automation capabilities explode but costs stay prohibitive, token efficiency finally makes autonomous systems pencil out financially

Google DeepMind dropped three new proprietary models on 21 July 2026. The company released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, positioning them as the most token-efficient models in the Gemini family to date.

Here's what "token-efficient" actually means: every time an AI model reads or writes text, it counts tokens. A token is roughly 4 characters of English text. If your AI agent processes a 10,000-word customer email thread, that's roughly 2,500 tokens. You pay per token. The fewer tokens required to do the job, the cheaper the task costs. Gemini 3.6 Flash is engineered to accomplish the same work with significantly fewer tokens than earlier models.

The three models target different use cases. Gemini 3.6 Flash is the main event — it's built for complex, multi-step reasoning tasks (the kind that take a while to think through). Gemini 3.5 Flash-Lite strips away some capability to prioritise raw speed and lower costs for simpler jobs. Gemini 3.5 Flash Cyber is hardened for security-sensitive workloads, which matters if you're processing customer financial data or confidential client records.

Google hasn't published exact pricing yet, but industry analysis suggests the cost per token will undercut current Claude and GPT-4 rates by 20–40%, and for certain task types, Gemini 3.6 Flash's token efficiency delivers that additional 65% saving on top.

The bigger signal: Gemini 3.5 Pro is on the way, and it's expected to be faster and more capable than Gemini 3 Pro while potentially costing less to run — a combination that would pressure the entire market to compete on efficiency.

Why does it matter for your business?

Imagine you run a 12-person accounting firm in Brisbane. Right now, you're manually reviewing expense reports from your clients' teams — spotting coding errors, flagging missing receipts, classifying transactions. It takes a junior accountant about 45 minutes per client per month. At $35/hour salary cost, that's $26.25 per client per month in pure labour. With 40 clients, you're spending $1,050 monthly on grunt work.

An AI agent could do this. It'd read the expenses, compare them to your firm's classification rules, flag anomalies, and produce a summary in 90 seconds. But six months ago, running that agent 24/7 cost $800–$1,200 monthly in API fees. The math didn't work. You'd save $1,050 in labour but spend $1,000 on tokens. Break even. Maybe lose money if the agent made mistakes.

Now, with Gemini 3.6 Flash, that agent costs $350–$500 monthly. The economics flip. You save $550 in net labour cost every month. Over a year, that's $6,600. Enough to hire someone part-time for something that actually generates revenue. Multiply that across your business — a few automated workflows — and you're looking at meaningful margin improvement.

According to analysis from AI infrastructure researchers, companies using token-efficient models like Gemini 3.6 Flash are reporting 40–55% reductions in AI infrastructure costs while maintaining output quality. For a business running multiple AI agents in parallel — one handling customer onboarding, another processing invoices, another analysing sales data — that difference scales fast.

There's another angle: speed. Gemini 3.5 Flash-Lite is optimized for latency. If you're using an AI agent for real-time customer service (answering questions in a chat, processing refund requests instantly), token efficiency directly translates to faster response times. Fewer tokens processed = quicker answers = happier customers.

---

65%
Average token cost reduction on long-horizon engineering tasks
Google DeepMind, 21 Jul 2026
$6,600
Annual labour savings for a 12-person firm automating expense review
247 AI News analysis
40–55%
Typical infrastructure cost reduction when switching to token-efficient models
AI infrastructure research community

---

Who should act on this right now?

  1. Boutique accounting and tax firms (10–30 people) — Gemini 3.6 Flash makes automating expense review, tax coding, and reconciliation profitable when it wasn't before.
  1. SaaS companies with customer support teams — AI agents handling tier-1 support or ticketing become viable at scale with these token costs; Gemini 3.5 Flash-Lite prioritises latency for real-time chat.
  1. Legal practices doing contract review — Long-horizon tasks like document analysis and contract classification are exactly what Gemini 3.6 Flash is engineered for; the 65% token saving is directly applicable.
  1. Digital agencies running client onboarding — Multi-step AI agentic workflows (collecting client info, setting up accounts, sending confirmations) are cheaper and faster now; automate your busy-work.
  1. Logistics and transport operators — Route optimisation, delivery scheduling, and invoice reconciliation are all-day AI agent tasks; token efficiency hits your operational bottom line hard.
  1. Medical and dental practices — Patient follow-up workflows, appointment scheduling, and claim processing are long-running, repetitive, and now feasible to automate cost-effectively with Gemini's security-hardened Cyber model.

Your 3 actions for this week

1
Check current spending on AI APIs.

Log into your Google Cloud Console, OpenAI dashboard, or Anthropic Claude portal and export your last 3 months of token usage. Google Cloud Console takes 10 minutes. Note the total monthly spend. You'll compare it against Gemini 3.6 Flash pricing once published — Google typically releases rates within 2 weeks of a model launch.

2
Identify one workflow you've wanted to automate but didn't because of cost.

It should be repetitive (done daily or weekly), well-defined (clear input and output), and long-running (takes an hour or more for a human). Write it down in 5

Source: VentureBeat · Verified and analysed by 247 AI News editorial team.