New · Launched July 21, 2026 · Gemini 3 series

Gemini 3.6 Flash — Free Trial17% Fewer Output Tokens

Gemini 3.6 Flash is Google DeepMind's workhorse agentic model, released July 21, 2026. According to the Artificial Analysis Index, it uses 17% fewer output tokens than Gemini 3.5 Flash, while bringing native computer use, a 1M-token context window, and lower output pricing to coding, research, and multi-step work. Start your free trial on Felo AI to evaluate it on your work.

Free trial on Felo AI — no credit card required

Agent economics

Evaluate cost and capability

83.0%

on OSWorld-Verified

$7.50

per 1M output tokens

Lower output pricing and stronger published agent benchmarks matter when work runs at volume.

83.0%

OSWorld-Verified

Agentic computer use benchmark

$7.50

Output Price

$7.50 / 1M tokens, down from $9.00

58.7%

SWE-Bench Pro

Agentic coding, up from 55.1%

1M

Context Window

Input tokens per request

The Headline Isn't a Benchmark. It's the Bill.

Agent costs are shaped by both token pricing and how reliably a model completes multi-step work. Gemini 3.6 Flash lowers output pricing from $9.00 to $7.50 per million tokens and, according to the Artificial Analysis Index, uses 17% fewer output tokens than Gemini 3.5 Flash.

17% fewer output tokens

Google DeepMind reports that Gemini 3.6 Flash uses 17% fewer output tokens than Gemini 3.5 Flash, according to the Artificial Analysis Index. Actual savings depend on the task and prompt.

Stronger published agent results

Google DeepMind reports gains over Gemini 3.5 Flash on DeepSWE, OSWorld-Verified, and long-context retrieval. These are useful signals when evaluating long-running agent tasks.

Same context, same modalities

You keep the 1M-token input window and native text, image, audio, and video input, so large and mixed-media tasks can stay in a single model workflow.

What Changed in Gemini 3.6 Flash

Built for the economics of running agents at scale, with official results that make cost and capability easier to evaluate together.

Pricing and Capability Together

Gemini 3.6 Flash lowers output pricing to $7.50 per million tokens from $9.00 on Gemini 3.5 Flash, while Google DeepMind reports 17% lower output token use according to the Artificial Analysis Index. Its published benchmark gains let teams evaluate cost and capability together instead of treating price as the only efficiency signal.

Computer Use as a Native Tool

Controlling a browser or desktop is now a built-in client-side tool through the Gemini API and Gemini Enterprise — no separate model or bolt-on framework. On OSWorld-Verified, agentic computer use reaches 83.0%, up from 78.4% on Gemini 3.5 Flash.

Stronger Agentic Coding

SWE-Bench Pro reaches 58.7% and DeepSWE v1.1 reaches 49.0%, compared with 55.1% and 37.0% for Gemini 3.5 Flash in Google's published comparison. Use the model card to assess fit for your own coding workflow.

1M Token Multimodal Context

Bring large codebases, long documents, and mixed media into one request. The 1M input window with 64K output tokens scored 91.8% on GDM-MRCR v2 at 128K in Google's published comparison.

Performance Benchmarks

Gemini 3.6 Flash vs the Frontier

Official Google DeepMind model card results. Gemini 3.6 Flash improves on Gemini 3.5 Flash across every row below; the comparison also shows where larger frontier models remain ahead.

Benchmark

Gemini 3.6 Flash

Gemini 3.5 Flash

Claude Sonnet 5

SWE-Bench Pro

58.7%

55.1%

63.2%

DeepSWE v1.1

49.0%

37.0%

54.0%

OSWorld-Verified

83.0%

78.4%

81.2%

CharXiv Reasoning (no tools)

85.2%

84.2%

77.0%

GDM-MRCR v2 (128K)

91.8%

77.3%

71.6%

Source: Gemini 3.6 Flash Model Card — Google DeepMind, July 2026.

Read the official Google DeepMind model card

Technical Specifications

Everything you need before integrating Gemini 3.6 Flash into your application.

Context Window

1,048,576 tokens input

65,536 tokens output

API Pricing

$1.50 / 1M input tokens

$7.50 / 1M output tokens

Output cost down from $9.00 on Gemini 3.5 Flash

Knowledge Cutoff

March 2026

Reasoning

A natively multimodal reasoning model designed for coding, computer use, and long-context tasks.

Tool Use & APIs

Function calling, structured output, and a new built-in client-side computer-use tool via the Gemini API and Gemini Enterprise.

Input Modalities

Text, images, audio, and video in; text out. Native multimodal, no preprocessing required.

Native Multimodal — One Model, Every Input Type

Gemini 3.6 Flash reads text, images, audio, and video natively. No separate pipelines, no stitching models together.

Text & PDF

Parses long documents and structured data in a single pass, handling complex tables and code without preprocessing.

Image & Charts

CharXiv Reasoning reaches 85.2% without tools in Google's published comparison, showing strong chart and visual reasoning.

Video Analysis

Analyze video alongside text and images, then ask grounded questions or request concise summaries of the material provided.

Audio Input

Understands speech, ambient sound, and multilingual conversation for transcription, translation, and voice-driven agents.

Who Uses Gemini 3.6 Flash

From individual developers to enterprise teams, Gemini 3.6 Flash fits wherever you run capable AI at volume.

Agentic Coding

Google's model card reports 58.7% on SWE-Bench Pro and 49.0% on DeepSWE v1.1. Evaluate it on your own repositories, review process, and coding-agent setup.

Document & Data Analysis

Use the 1M-token context window to work across lengthy documents, structured data, and charts in a single request when that context matters.

Design & Product Tools

Combine visual input and written instructions to extract structured information, draft alternatives, or prepare material for review.

Computer-Use Agents

With computer use as a native tool and OSWorld-Verified at 83.0%, agents can drive browsers and apps directly — booking, filling forms, and navigating UIs end to end.

Long-Horizon Workflows

For multi-step work, compare model quality, latency, and token usage on representative tasks before choosing a production setup.

High-Volume Pipelines

For high-volume document and extraction workflows, measure actual request volume and output length to estimate API costs before rollout.

Two Ways to Use Gemini 3.6 Flash on Felo

Felo AI Search

Open Felo AI Search and select the Gemini 3.6 Flash model. Ask questions, search the web with AI, and get cited answers — powered by Google's newest workhorse model.

Open Felo AI Search

Compare Answers in Felo Search

Open Felo AI Search, choose Gemini 3.6 Flash where available, and compare its answers with other supported models on your own prompts.

Open Felo AI Search

Frequently Asked Questions

Yes. Felo AI offers a free trial of Gemini 3.6 Flash. Sign up for an account to get started — no credit card required.

Start Your Gemini 3.6 Flash Free Trial

Google DeepMind's workhorse model, published July 2026. Open Felo AI and evaluate its coding, computer-use, and long-context capabilities on your own work.

Open Gemini 3.6 Flash on Felo

Free trial — no credit card required