Search Ctrl+K
Change language Switch ThemeSign In
Curated Daily BriefWeekly PicksTopics SettingsHelp CenterCollapse
Narrow Mode
Google Releases Three New Gemini Models Including 3.6 Flash, with Improved Performance and Cost-Effectiveness
Google Releases Three New Gemini Models Including 3.6 Flash, with Improved Performance and Cost-Effectiveness
 宝玉@dotey
Google 今天一口气发布了三款 Gemini 新模型:Gemini 3.6 Flash、3.5 Flash Lite 和 3.5 Flash Cyber。主角是 3.6 Flash,一个在多个维度上都比前代 3.5 Flash 更好、同时还更便宜的模型。
先看最直观的变化。3.6 Flash 的输出定价从 3.5 Flash 的每百万 token $9 降到了 $7.50,输入价格不变($1.50)。
除了便宜还省 token。根据 Google 的数据,3.6 Flash 在 Artificial Analysis Index 上平均少用 17% 的输出 token,在某些 DeepSWE 编码测试中,token 消耗最多降低 65%。Artificial Analysis 的实测也印证了这一点:跑同一套智能评测,3.5 Flash 花了 $1041,3.6 Flash 只花了 $727,账单直接少了三成。
速度也有明显提升。3.6 Flash 的生成速度达到 304 tokens/s,3.5 Flash 是 165 tokens/s,接近翻倍。对于需要多步推理、反复调用工具的 Agent 工作流来说,每一步快一倍,最终的完成时间差距会被放大。
智能水平方面,3.6 Flash 在 Artificial Analysis Intelligence Index 上得分 50,和 3.5 Flash 持平,但用更少的 token 达到了同样的分数。在知识工作基准 GDPval-AA 上,3.6 Flash 拿到 1421 分,3.5 Flash 是 1349。Computer Use 能力在 OSWorld-Verified 上从 78.4% 提升到 83%。知识截止日期也从 2025 年 1 月跳到了 2026 年 3 月,意味着模型终于知道过去一年多的事了。
放到竞品里看,3.6 Flash 在 OSWorld 和两项长上下文测试上领先,但 GPT 5.6 Luna 在 DeepSWE 和 Terminal-bench 上更强,Grok 4.5 在 SWE-Bench Pro 上领先,Claude Sonnet 5 在 MLE-Bench 和 GDPVal-AA v2 上得分最高。四家模型在多数基准上差距不大,实际选择更多取决于具体任务和价格。
3.6 Flash 已经可以在 GitHub Copilot 中使用,同时上线了 Gemini 应用、Google AI Studio 和 Google Antigravity。Antigravity 是 Google 今年 5 月在 I/O 大会推出的 Agent 开发平台,定位是以 Agent 编排为核心的独立桌面应用,有点像 Google 版的 Claude Code + Cursor 的合体。
另外两个模型定位不同。3.5 Flash Lite 走极致低价路线,定价 $0.30/$2.50(每百万 token 输入/输出),适合搜索、文档处理等高吞吐量场景。3.5 Flash Cyber 专注网络安全,目前仅对政府和受信任合作伙伴开放。
不过这次发布 Gemini 3.5 Pro 缺席。Google 在 I/O 上承诺的旗舰模型 3.5 Pro 至今没有公开发布,Bloomberg 此前报道,延期原因是模型在内部编码基准上表现不达预期,6 月底用新数据重新训练后结果仍然不理想。目前 3.5 Pro 只在少数合作伙伴和美国政府那里做测试,公开日期未知。
上次 3.5 Pro 跳票时,Google 推出了 Gemini 3.5 Flash,结果这个 Flash 模型反而在多项编码和 Agent 基准上打败了 Gemini 3.1 Pro。Gemini 3.6 Flash 看起来也是类似:不需要和 GPT-5.6 或 Claude 争排行榜顶端,先把更快更便宜的生态位占住。
Google 同时确认,下一代 Gemini 4 已经启动了迄今最大规模的预训练。Show More
@Google · 5h ago
Today we’re expanding the Gemini family with three new models built to be faster, more token efficient, and reliable at scale.
Meet the new Gemini models ↓Show More
00:10
301
537
4,031
400.4K
Jul 21, 2026, 5:25 PM View on X
17 Replies
3 Retweets
35 Likes
8,993 Views  宝玉 @dotey
Follow
One Sentence Summary
Google released three new Gemini models today: 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber, with 3.6 Flash offering significant improvements in price, speed, and token efficiency.
Summary
Google released three new Gemini models today: Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber. The output price for 3.6 Flash dropped from $9 to $7.50 per million tokens, while the input price remains at $1.50. On the Artificial Analysis Index, it uses an average of 17% fewer output tokens, with token consumption reduced by up to 65% in certain DeepSWE coding tests. Real-world tests confirmed this: running the same intelligence benchmark, 3.5 Flash cost $1,041, while 3.6 Flash cost only $727, a roughly 30% reduction. Generation speed increased to 304 tokens/s, nearly double the 165 tokens/s of 3.5 Flash. For Agent workflows requiring multi-step reasoning and repeated tool calls, doubling the speed at each step amplifies the overall completion time advantage. In terms of intelligence, 3.6 Flash scored 50 on the Artificial Analysis Intelligence Index (matching 3.5 Flash), but achieved the same score with fewer tokens. On the knowledge work benchmark GDPval-AA, it scored 1,421 (vs. 1,349 for 3.5 Flash). Its Computer Use capability on OSWorld-Verified improved from 78.4% to 83%. The knowledge cutoff date also jumped from January 2025 to March 2026, meaning the model is now aware of events from the past year. Compared to competitors, 3.6 Flash leads in OSWorld and two long-context tests, but GPT-5.6 Luna is stronger in DeepSWE and Terminal-bench, Grok 4.5 leads in SWE-Bench Pro, and Claude Sonnet 5 scores highest on MLE-Bench and GDPVal-AA v2. The four models show little difference on most benchmarks, so practical choice depends more on specific tasks and pricing. 3.6 Flash is already available in GitHub Copilot, and has also launched on the Gemini app, Google AI Studio, and Google Antigravity. Antigravity is an Agent development platform unveiled by Google at I/O in May this year, positioned as an independent desktop application centered on Agent orchestration, somewhat like a combination of Claude Code and Cursor from Google. The other two models have different positioning. 3.5 Flash Lite follows an ultra-low-cost route, priced at $0.30/$2.50 (input/output per million tokens), suitable for high-throughput scenarios like search and document processing. 3.5 Flash Cyber focuses on cybersecurity and is currently only open to government and trusted partners. Notably, Gemini 3.5 Pro is absent from this release. The flagship model 3.5 Pro, promised by Google at I/O, has yet to be publicly released. Bloomberg previously reported the delay was due to the model's underperformance on internal coding benchmarks, and results remained unsatisfactory after retraining with new data at the end of June. Currently, 3.5 Pro is only being tested with a few partners and the U.S. government, with no public release date. When 3.5 Pro was previously delayed, Google launched Gemini 3.5 Flash, which ironically outperformed Gemini 3.1 Pro on multiple coding and Agent benchmarks. Gemini 3.6 Flash seems to follow a similar pattern: it doesn't need to compete for the top spot on leaderboards against GPT-5.6 or Claude, but instead secures the niche of being faster and cheaper. Google also confirmed that the next-generation Gemini 4 has already begun its largest-scale pre-training to date.
AI Screening
87
Influence Score 15
Published Today
Language
Chinese
Tags
Gemini 3.6 Flash
Gemini 3.5 Flash Lite
Gemini 3.5 Flash Cyber
Gemini 4
AI Models
Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free HomeDiscoverSettings