Gemini 3.6 Flash: When to Use It Over Claude and GPT
Gemini 3.6 Flash is Google's new low cost Flash model, and on Google's own benchmark it beats Claude Sonnet 5, GPT-5.6 Luna, and Grok 4.5 at computer use, reading charts, long video, and long context. It loses at deep coding. So the honest answer to when you should reach for it is short: use Gemini 3.6 Flash for eyes and memory work, the documents and the screens, and keep deep coding on the flagship models. This post walks through the numbers Google published, where the wins hold up, where the losses are, and how I use it in my own work every week.
The AI space moves weekly, so every price and benchmark below is current at the time of writing and linked to its source.
Key Takeaways
- Gemini 3.6 Flash beats Claude Sonnet 5, GPT-5.6 Luna, and Grok 4.5 at computer use, chart reading, long video, and long context on Google's own benchmark.
- It loses at deep coding. On SWE-Bench Pro it scores 58.7, behind Grok 4.5 at 64.7, Claude Sonnet 5 at 63.2, and GPT-5.6 Luna at 62.7.
- At the time of writing it costs $1.50 per million input tokens and $7.50 per million output tokens, about half Claude Sonnet 5's price, though GPT-5.6 Luna is cheaper still.
- It is fast. Artificial Analysis clocks it around 305 output tokens per second, ahead of GPT-5.6 Luna and well ahead of Claude Sonnet 5.
- The comparison is against mid tier models, not each brand's flagship, so a top model like Opus 4.8 or GPT-5.6 Sol may still win these rows.
What Is Gemini 3.6 Flash?
Gemini 3.6 Flash is a low cost, high speed model in Google's Gemini family, built for agentic and multimodal tasks. Multimodal means it takes several forms of media at once, text, images, video, and audio, inside a single pipeline, rather than bolting them together. It carries a 1M token context window, up to 64k tokens of output, and Google's built in tool suite, including Computer Use.
There is a version quirk worth knowing. Flash is now on 3.6 while Google's Pro line is still on 3.1, so the cheap tier has leapfrogged the flagship in version number. Google released it alongside a smaller sibling, Gemini 3.5 Flash-Lite.

At the time of writing, Google lists Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. That output price is lower than the previous Gemini 3.5 Flash, which ran $9.00 per million output tokens on the same input price.
Where Gemini 3.6 Flash Beats Claude, GPT, and Grok
Gemini 3.6 Flash wins the categories that are about seeing and remembering, not writing code. On Google's own comparison, it tops Claude Sonnet 5, GPT-5.6 Luna, and Grok 4.5 at computer use, chart reading, long video, and long context.

Here are the headline wins against the model most people compare it to, Claude Sonnet 5. According to Google's evaluation methodology, the numbers are:
| Benchmark (what it measures) | Gemini 3.6 Flash | Claude Sonnet 5 |
|---|---|---|
| OSWorld, computer use | 83.0 | 81.2 |
| CharXiv, reading charts | 85.2 | 77.0 |
| MRCR 128k, long context | 91.8 | 71.6 |
The long context result is the one I would not skip past. On the MRCR test at a full 1M tokens, Gemini 3.6 Flash scores 54.0, where GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5 post no score at that length at all. Hand it hundreds of pages, or a folder of screenshots, and it will find the answer while the others are still declining to try.
One honest caveat sits under all of this. These are Google's own benchmarks, run by the vendor against its rivals, so treat the wins as directional. The reason they still read as credible is that Google published a chart where its own model loses three coding rows outright. A pure marketing chart does not do that.
Where Gemini 3.6 Flash Loses: Deep Coding
Gemini 3.6 Flash loses at deep coding, and it is not close on the hardest rows. This is the one job to keep on the flagship models.

On SWE-Bench Pro, a broad agentic coding test, Gemini 3.6 Flash scores 58.7, behind Grok 4.5 at 64.7, Claude Sonnet 5 at 63.2, and GPT-5.6 Luna at 62.7. On DeepSWE it scores 49 against GPT-5.6 Luna at 67. On knowledge work, measured as an Elo score, it lands at 1421 while Claude Sonnet 5 reaches 1607. If code generation is the core of your work, that gap is the whole story.
One more thing to keep straight. This chart compares Gemini 3.6 Flash against the mid tier models, Claude Sonnet 5 and GPT-5.6 Luna, not each brand's premier model. A flagship like Opus 4.8 or GPT-5.6 Sol may beat Gemini even on the rows it wins here. So a fair reading of the win is "beats Claude Sonnet 5 and GPT-5.6 Luna," not "beats Claude and GPT outright." The same logic applies inside one brand, which is why I wrote up Sonnet vs Opus and when the cheaper one is the right call.
Speed and Price: Cheaper Than Claude, Faster Than Both
Gemini 3.6 Flash is both cheaper than Claude and faster than Claude and GPT. That combination is the real reason to reach for it on high volume work.

On speed, Artificial Analysis places Gemini 3.6 Flash alone in the top right quadrant at around 305 output tokens per second. GPT-5.6 Luna is next at about 190, with Claude Sonnet 5 near 85 and Grok 4.5 slower still. On price, at the time of writing Gemini 3.6 Flash runs about half the cost of Claude Sonnet 5, which is $3.00 input and $15.00 output per million tokens at full price.
The price claim holds against Claude, and only against Claude. GPT-5.6 Luna is actually cheaper, at $1.00 input and $6.00 output per million tokens. So the accurate line is cheaper than Claude, faster than both.
The Flagship Google Is Still Holding Back
Gemini 3.6 Flash arrived while Google's flagship stayed home. Google reportedly delayed its top tier Gemini 3.5 Pro, according to a report from Android Headlines that cited code generation issues. I say reportedly because that is a secondhand report, not Google's own statement, so treat it as a signal rather than a fact.
The read is straightforward. Google released upgrades to the cheap and free tiers people already use, and held back the premier model many were waiting for. That is worth knowing before you assume the whole Gemini line just moved forward. The Flash tier did. The flagship, at the time of writing, has not.
How I Actually Use Gemini 3.6 Flash
I reach for Gemini as a verification layer whenever I need to process video or audio, because handling both natively is where it is genuinely strong. When I have a stack of PDFs with heavy context to sort through, that is a Gemini job too. It is good enough at that work, and it costs a fraction of what a flagship would.
There is a more concrete example sitting in this very post. That news card above, the one about the delayed flagship, is from AI Matters, a news site I run. The synthesis behind each story on it runs on Gemini. It is cheap enough to run in production and the output holds up, which is exactly the trade this whole post is about: match the model to the job, and a cheaper model wins more often than people expect.
When to Use Gemini 3.6 Flash vs Claude or GPT
The decision comes down to the shape of the task. Here is the rule I follow.
| The task | Reach for | Why |
|---|---|---|
| Computer use, charts, long video | Gemini 3.6 Flash | Leads the benchmark and costs far less |
| Huge documents, long context | Gemini 3.6 Flash | 1M context window, tops the long context test |
| Deep coding | Claude Sonnet 5 or GPT-5.6 Luna | Gemini loses the coding rows |
| Knowledge work quality | Claude Sonnet 5 | Higher score on the knowledge work test |
I put the whole decision on one page, with the benchmark numbers, the price, and the speed, plus the reminder that these are Google's own figures. Enter your email in the card below and I'll send it over, so it is on hand right before your next task, when you are not sure which model fits.

If deep coding is where you spend most of your time, you can still cut the bill without dropping quality. The advisor setup gets you most of a top model's quality for a fraction of the price, by having a cheap model do the work and a smarter one weigh in only at the hard calls.

