We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.
No reviews yetBe the first to leave a review for Gemini 3.6 Flash Family
The positioning around “efficiency, latency, and reliability for agents at scale” feels incomplete. Right now every claim is qualitative no hard numbers, no latency charts, no reliability benchmarks, no failure‑mode transparency.
Developers building multi‑step agents don’t just need fast models; they need predictable tool‑use behaviour across long chains. Without published consistency metrics, uptime guarantees, or real‑world agent stress tests, it’s impossible to evaluate whether 3.6 Flash actually solves the reliability gaps that break production agents.
Also: Flash‑Lite and Flash‑Cyber are introduced as if they’re cleanly differentiated, but the docs force us to dig through scattered pages to compare cost, reasoning quality, and safety behaviour. A unified dashboard isn’t a “nice to have” it’s essential.
Right now this launch reads more like marketing than engineering. If Google wants developers to commit workloads, we need numbers, not adjectives.
Report
Congrats@sundar_pichaion the launch. Exciting to see 3.5 Flash Cyber. Since it's gated to governments/trusted partners for now; is wider access on the roadmap, or is this meant to stay a permanent limited-access tool given the dual-use risk?
Report
BTW, I'm also curious how you balance "stronger safety guardrails" with "fewer refusals" cos those feel like they'd pull in opposite directions. How do you know when you've got that balance right?
Report
Been testing flash models quite heavily lately. The lite variant is interesting for high volume use cases where cost matters more than raw performance. How does 3.6 compare to 3.5 on reasoning tasks?
Report
Smart that the push is on the Flash tier rather than another flagship — for agents the cost and latency per call matter way more than a few extra points on a frontier benchmark, since a single run fires off so many calls. The jumps on OSWorld (computer use) and long-horizon SWE are the ones I'd actually feel day to day.
Curious what "3.5 Flash Cyber" is tuned for — is that a security/red-team focused variant, or something else?
Report
A unified dashboard to compare latency, cost, and quality across the three Flash variants side by side would be huge. Right now I have to dig through docs and run my own benchmarks to figure out whether Flash Lite or Flash Cyber fits a given workload, and a built-in comparison view with sample prompts would save a ton of evaluation time before committing to a model.
Report
Love the focus on efficiency for agent workflows. One thing that would help me as a developer is a built-in token usage dashboard per model variant, so I can compare cost and latency between Flash, Flash-Lite, and Flash Cyber in real time during testing without needing separate tooling. Would save a lot of guesswork when picking the right model for each part of a pipeline.
Congrats@sundar_pichaion the launch. Exciting to see 3.5 Flash Cyber. Since it's gated to governments/trusted partners for now; is wider access on the roadmap, or is this meant to stay a permanent limited-access tool given the dual-use risk?
BTW, I'm also curious how you balance "stronger safety guardrails" with "fewer refusals" cos those feel like they'd pull in opposite directions. How do you know when you've got that balance right?
Been testing flash models quite heavily lately. The lite variant is interesting for high volume use cases where cost matters more than raw performance. How does 3.6 compare to 3.5 on reasoning tasks?
Smart that the push is on the Flash tier rather than another flagship — for agents the cost and latency per call matter way more than a few extra points on a frontier benchmark, since a single run fires off so many calls. The jumps on OSWorld (computer use) and long-horizon SWE are the ones I'd actually feel day to day.
Curious what "3.5 Flash Cyber" is tuned for — is that a security/red-team focused variant, or something else?
A unified dashboard to compare latency, cost, and quality across the three Flash variants side by side would be huge. Right now I have to dig through docs and run my own benchmarks to figure out whether Flash Lite or Flash Cyber fits a given workload, and a built-in comparison view with sample prompts would save a ton of evaluation time before committing to a model.
Love the focus on efficiency for agent workflows. One thing that would help me as a developer is a built-in token usage dashboard per model variant, so I can compare cost and latency between Flash, Flash-Lite, and Flash Cyber in real time during testing without needing separate tooling. Would save a lot of guesswork when picking the right model for each part of a pipeline.