Gemini 3.8 Flash: 4 Jobs It Is Worth Switching For
Gemini 3.8 Flash is Google's new Flash model, and on Google's own benchmark table it lands within a rounding error of Claude Opus 5 on long horizon coding while costing $0.75 per million input tokens on the API against Opus 5's $5.00. It also takes the top score on financial analyst work and on complex legal workflows. On the hardest general agent tasks it is a long way back. So the honest answer to whether you should switch is: switch the coding agents, the document work, and anything multimodal running in a loop, keep the open ended computer use on a flagship model.
The AI space moves weekly, so every price and benchmark below is current at the time of writing and linked to its source. The benchmark numbers are Google's own published set, which makes them a starting point for your own evaluation rather than a settled ranking.
Key Takeaways
- Coding agents that run long are the clearest switch. On DeepSWE v1.1, long horizon software engineering, it scores 73.7% against Claude Opus 5 at 74.0% and GPT-5.6 Sol at 72.7%.
- Finance and legal agents are the second. It takes the top score on Vals Finance Agent v2 at 61.4% and on Harvey's Legal Agent Benchmark at 10.0%, where the leading score still means 1 task passed in 10.
- Keep open ended computer use on a flagship. On OSWorld-2.0 it scores 59.0% against Claude Opus 5 at 75.4%, and on the hardest terminal tasks it is further back still.
- On the API, at the time of writing, it costs $0.75 per million input tokens and $3.75 per million output tokens. That introductory rate ends on December 31, 2026. After that it doubles to $1.50 and $7.50.
- On a Google AI Pro or Ultra subscription it arrives with no pricing decision to make. The per token rates matter when you are building something that runs on its own.
Short on time? Jump to 4 jobs worth switching for.
What Gemini 3.8 Flash Is
Gemini 3.8 Flash is the third Flash release in roughly 6 weeks, and Google positions it as its "most intelligent workhorse model", built for long horizon software engineering, autonomous agents, and multi step reasoning in specialist fields. The pitch is that it holds the speed and the price of Gemini 3.7 Flash while scoring much closer to the frontier models.
What does it cost?
On the API, $0.75 per million input tokens and $3.75 per million output tokens, at the time of writing. That is an introductory rate that expires on December 31, 2026. From January 1, 2027 it becomes $1.50 and $7.50, which is still the same price as Gemini 3.7 Flash.
For scale, Claude Opus 5 runs at $5.00 and $25.00, Claude Sonnet 5 at $2.00 and $10.00, GPT-5.6 Sol at $4.00 and $20.00.
How big is the context window?
1,048,576 input tokens, with a maximum of 65,536 output tokens per response, according to the model documentation. It takes text, images, video, audio, and PDFs as input, so a single call can hold a long document set and a video without a separate pipeline to prepare them.
What effort levels does it support?
Low, medium, and high, which set how hard the model thinks before it answers. There is one detail worth knowing before you wire it up: minimal is not supported and returns an error. If you are moving code across from a model that accepted it, that call will fail rather than fall back.
Where Gemini 3.8 Flash Matches the Frontier Models
3 areas, and it is worth being precise about what "matches" means in each, because some of these rows are ties rather than wins.
Long horizon software engineering
On DeepSWE v1.1, the benchmark for work that runs across many steps and files, Gemini 3.8 Flash scores 73.7%. Claude Opus 5 scores 74.0%. GPT-5.6 Sol scores 72.7%. Those 3 models are separated by 1.3 percentage points, and Gemini 3.8 Flash costs a sixth of what Claude Opus 5 does per input token.
The jump inside the Flash line is the bigger story. Gemini 3.7 Flash scores 65.3% on the same benchmark, so this is an 8.4 point move in one release.
Finance and legal agent work
This is where the top score is unambiguous. On Vals Finance Agent v2, financial analyst tasks, Gemini 3.8 Flash scores 61.4%, ahead of Claude Opus 5 at 58.6% and GPT-5.6 Sol at 53.8%.
On Harvey's Legal Agent Benchmark it scores 10.0%, ahead of Claude Opus 5 at 6.7% and GPT-5.6 Sol at 2.5%. Read that number twice. The best available model on complex legal workflows completes 1 task in 10. If you are about to point an agent at professional work with a deadline attached, that figure is the one to plan around, and the fact that it is the leading score is exactly why it matters.
Charts, long video, and expert reasoning
Gemini 3.8 Flash leads on LVBench, understanding long video, at 87.8% in Google's agentic setting and 87.1% static, against 82.1% for GPT-5.6 Sol and 75.4% for Claude Opus 5. The lead holds on the more conservative of its own two figures, and long video is not something you can work around with a bigger prompt. On CharXiv, reading information out of complex charts with no tools, it is level with the field rather than ahead of it: 86.2% against 85.9% and 85.8%.
On HLE-Verified, multidisciplinary expert reasoning, it scores 54.9% against GPT-5.6 Sol at 54.5% and Claude Opus 5 at 54.4%. Half a point across 3 models is a tie, not a win, and a tie at this price is the useful finding.
Here is the full table Google published, covering 16 benchmarks across 6 models:
Where Gemini 3.8 Flash Is Not Close
The same table carries the losses, and some of them are wide.
General agent capability
Terminal-bench 2.1 and Terminal-bench 4.0 sound like one benchmark at two version numbers, and the scores say otherwise. On 2.1, mid difficulty terminal coding, Gemini 3.8 Flash scores 89.4% and Claude Opus 5 scores 89.1%, so they are level. On 4.0, a much harder set of tasks, Gemini 3.8 Flash scores 19.1% and Claude Opus 5 scores 51.8%.
That is the single most important pair of numbers in the table. Level on mid difficulty terminal work, 32.7 points back on the hardest tasks in the same family.
Computer use
On OSWorld-2.0, driving a computer directly, Gemini 3.8 Flash scores 59.0% against Claude Opus 5 at 75.4% and GPT-5.6 Sol at 62.6%. It is ahead of Claude Sonnet 5 at 42.6%, so it is not weak in absolute terms. It is simply not the model to hand a screen to when the task is not scripted.
Knowledge work
On GDPVal-AA v2, general professional output, the scoring is a head to head rating rather than a percentage, so the gap matters more than the number. Gemini 3.8 Flash sits at 1545 against Claude Opus 5 at 1824 and GPT-5.6 Sol at 1710. For everyday knowledge work the flagships are still well ahead.
4 Jobs Worth Switching to Gemini 3.8 Flash For
Reach for it when the work looks like this.
- Coding agents that run long and run often. Multi file, multi step engineering work where the task is specified and the model has to stay on it. The DeepSWE gap to Opus 5 is 0.3 points, and the price gap is roughly 6 times on input while the introductory rate holds, a little over 3 times after it.
- Document and analysis agents in finance or legal. It holds the top score on both, ahead of Claude Opus 5 rather than level with it, and both are workloads people are actively building agents for right now.
- Anything multimodal in a loop. Charts, long video, PDFs, audio, all in one call with a 1 million token window. If your pipeline currently splits a video and a document across 2 models, this collapses it into 1.
- High volume production agents, but only after you measure. It scores 59 on the Artificial Analysis Intelligence Index at high reasoning, at $0.58 per task on that index. At high effort it spends more tokens per task than 3.7 Flash does, so run your own workload both ways before you commit.
When to keep the work somewhere else
Keep it on Claude Opus 5 or GPT-5.6 Sol when the agent has to work out what to do rather than execute a plan: open ended computer use, general knowledge work, or a long tail of tasks harder than anything you would think to write down. Choosing between the flagship tiers is its own decision, and I have written up effort settings separately.
There is a second way to not switch, and it is not upward. In the launch announcement, the recommendation for efficiency first work is to drop to a lower effort level or stay on Gemini 3.7 Flash, because 3.8 Flash "works harder", running extra reasoning steps and calling tools repeatedly, and spending more tokens to do it. A model that is cheap per token is not automatically cheap per task, and Google says so plainly.
Google Built a Second Model: Gemini 3.8 Flash Cyber
Launched the same day and almost entirely missing from the coverage, Gemini 3.8 Flash Cyber is a security specialist built for finding vulnerabilities and writing patches for them.
What Cyber is for
It ships with a more permissive set of safety mitigations for security work, which is what lets it examine and repair vulnerable code that the standard model would decline to handle. On Google's own numbers, it produced 2.6 times more correct patches to Chrome vulnerabilities than the commercial models the Chrome security team compared it against, and it found a critical vulnerability for Google Cloud's research team in under 2 hours, on a class of problem where discovery normally takes months. These are a vendor's figures for its own model on its own codebase, which is worth holding in mind.
On CWE-Bench, finding and fixing real code vulnerabilities, it gets 47.2% right on the first attempt against 47.8% for a leading frontier model. By the same standard applied above that is level, not ahead.
Who can get it
Most people cannot. It is available only through a new application based Fairwind Program, aimed at government authorities, critical infrastructure operators, and software maintainers. If you maintain something widely depended on, it is worth an application. Otherwise this is a signal about where security tooling is heading rather than a tool you can pick up this week.
What This Means If You Pay a Subscription, Not an API Bill
Every price above is an API rate, and most people reading this pay a flat monthly fee instead. For Gemini specifically, 3.8 Flash is already in the Gemini app, in AI Mode in Search, and in Google Sheets for Google AI Pro and Ultra subscribers, so on that plan the model arrives without a pricing decision to make.
The per token numbers still matter to you in one case: when you are building something that runs on its own, repeatedly, and the bill scales with how much it does. That is the situation the 4 jobs above describe. If you are working interactively inside a subscription, the benchmark rows are the part to read and the price column is context rather than a cost you will meet. The same split applies to model tier pricing elsewhere.
For the previous release in this line and how its trade offs differed, the Gemini 3.6 Flash numbers are worth a look, mostly because that comparison ran against mid tier models while this one includes both flagships. And if you want to work through which model belongs on which job in your own setup rather than reading a table, that is most of what I cover in the live workshop.

