Skip to content

Gemini 3 Deep Think vs Flash vs Pro: Thinking Levels Explained

Deep Think lives only in the Gemini app (Ultra, and reportedly AI Pro). Flash and Pro are also API models, and thinking_level decides how hard they reason.

•••15 min read•AI Model Comparison
Gemini 3 Deep Think vs Flash vs Pro complete comparison guide

"Deep Think", "Flash", and "Pro" are not three models on one menu. Deep Think is a reasoning mode inside the Gemini app, available to Google AI Ultra subscribers (and, according to reports, to AI Pro subscribers from October 2026), and it takes minutes per answer. Flash and Pro exist in two places at once: as options in the app's model picker and as API models (gemini-3.8-flash, gemini-3.5-flash, gemini-3.1-pro-preview) whose reasoning depth you set per request with thinking_level. There is no Deep Think model ID in the API models list, the pricing page, or the thinking documentation as of October 9, 2026.

If you are in the app and torn between Thinking and Pro, the short answer is: Pro is the Gemini 3.1 Pro model, and the effort level you pick decides how long it thinks. If you are in the API and want Flash to stop thinking as much as possible, set thinking_level to minimal on gemini-3.5-flash or gemini-3.6-flash; the newest gemini-3.8-flash only goes down to low. The sections below give the full per-model table, the plan changes that start today, and what Deep Think actually gets you.

Deep Think vs Flash vs Pro at a glance: one table across app and API

Everything else in this guide expands on this one table. Prices are Standard tier, per 1M tokens, from the Gemini API pricing page; levels and defaults are from the thinking documentation; plan access is from Google's model access help page, all as of October 9, 2026.

Name you seeWhat it isAPI model IDthinking_level values (default)API price in / outWho can use it in the app
Flash-LiteSmallest Flashgemini-3.5-flash-liteminimal / low / medium / high (minimal)listed on the pricing pageEveryone, including no subscription
Flash (shown as "Fast" and "Thinking" in late 2025)The default workhorse modelgemini-3.8-flash (newest stable), gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash (legacy)3.8 and 3.7: low / medium / high (medium); 3.6 and 3.5: minimal / low / medium / high (medium)3.5 Flash: $1.50 / $9.00, free tier availableAI Plus, AI Pro, AI Ultra
ProThe top-end modelgemini-3.1-pro-previewlow / medium / high (high); no minimal$2.00 / $12.00 up to 200k prompt tokens; $4.00 / $18.00 above; no free tierAI Pro, AI Ultra
Deep ThinkA reasoning mode on top of Pro, answers in minutesnone documentednot applicableno API price; included in the subscriptionAI Ultra; AI Pro according to October 2026 reports
Gemini app picker labels mapped to Gemini API model IDs and thinking levels

Two things in this table surprise most people. First, "Gemini 3 Pro" is no longer an API model: gemini-3-pro-preview is listed as shut down on the models page, and the Pro lane is gemini-3.1-pro-preview. Second, the cheapest "Flash" price you may remember, $0.50 input and $3.00 output, belongs to gemini-3-flash-preview, which is still available but is a preview model from December 2025; the current 3.5 Flash costs $1.50 and $9.00.

Gemini Thinking vs Pro in the app: what Fast, Thinking, and Pro mean

In the Gemini app, "Thinking" has meant two different things, which is why the question keeps coming up. According to the Gemini app release notes:

  • November 18, 2025: Gemini 3 Pro arrived for everyone in the app, and you reached it by "selecting 'Thinking' in the model drop down". For a month, Thinking was Pro.
  • December 17, 2025: Gemini 3 Flash became the default. The picker changed to "Fast" for quick answers and "Thinking" for harder problems, both running on Gemini 3 Flash, while Gemini 3 Pro moved to its own "Pro" entry. From this point on, Thinking was Flash with more reasoning, and Pro was the bigger model.
  • February 19, 2026: Gemini 3.1 Pro rolled out behind the "Pro" entry, with higher limits for AI Pro and AI Ultra.
  • May 19, 2026 and July 21, 2026: 3.5 Flash and then 3.6 Flash became selectable in the model drop-down.

So "Flash vs Thinking vs Pro" from the December 2025 picker maps to: Fast = Gemini 3 Flash answering quickly, Thinking = Gemini 3 Flash with extended reasoning, Pro = Gemini 3 Pro (later 3.1 Pro). Any "Thinking vs Pro" comparison from that period is comparing Flash with extended reasoning against the Pro model, not two tiers of one model.

The current app no longer uses the Fast/Thinking split as the main choice. Google's model access page describes a model list of Flash-Lite, Flash, and Pro, and for each available model you can select an effort level of low, medium, or high. That effort level is the consumer-side counterpart of the API's thinking_level. "Extended thinking and Deep Think" are listed as premium features that use more of your usage limit, so a Pro answer at high effort uses up your plan's limit faster than a Flash answer at low effort.

Thinking or Pro for your question? Pick Pro when the task is long multi-step reasoning, large documents, or code you will actually run; pick Flash (the former Thinking) at medium or high effort when you want a considered answer quickly and your plan's Pro quota is the scarce resource. For everyday questions, Flash at low effort or Flash-Lite is the right default, and that is also the only thing a non-subscriber gets from today.

Which Gemini models each plan gets from October 9, 2026

Google's model access page says the changes "will start to take effect for users without an AI subscription on October 9th", and that AI Plus subscribers will get an email with their date. The table on that page, plus the AI Ultra benefits page and a 9to5Google report of October 3, 2026 for the parts Google has not written down, gives this picture:

PlanModels in the Gemini appDeep Think
No subscriptionFlash-Lite only (Flash and Pro removed from October 9, 2026)No
Google AI Plus ($4.99/month per the report)Flash-Lite and Flash; Pro removed on a date sent by emailNo
Google AI Pro ($19.99/month per the report)Flash-Lite, Flash, ProAccording to the report, gains the Deep Think option; Google's page does not list it
Google AI Ultra (starting at $100/month per the May 19, 2026 release note)Flash-Lite, Flash, Pro, with 5x or 20x the usage quota of AI ProYes, "3 Pro Deep Think" for eligible Ultra plans

Google's own page does not say which Flash version each plan gets; the report names 3.6 Flash and 3.5 Flash-Lite as the models involved. If you are on AI Plus and rely on Pro, treat the email date as your deadline rather than today.

Gemini 3 thinking levels in the API: minimal, low, medium, high by model

Every current Gemini text model thinks by default; thinking_level controls how much. The values are minimal, low, medium, and high, but not every model accepts every value, and the default differs by model. This table is the current one from the thinking documentation as of October 9, 2026:

ModelSupported thinking_levelDefault
gemini-3.8-flash, gemini-3.7-flashlow, medium, highmedium
gemini-3.6-flash, gemini-3.5-flashminimal, low, medium, highmedium
gemini-3.5-flash-liteminimal, low, medium, highminimal
gemini-3.1-pro-previewlow, medium, highhigh
gemini-3-flash-previewminimal, low, medium, highhigh
gemini-2.5-pro, gemini-2.5-flashlow, medium, highthinking on
gemini-2.5-flash-litelow, medium, highthinking off

Three rules from the Gemini 3 developer guide change how you should read that table:

  • Levels are "relative allowances for thinking rather than strict token guarantees". high on Flash-Lite is not the same number of thought tokens as high on Pro.
  • minimal "does not guarantee that thinking is off". It is the lowest allowance, not a switch. The 2.5 Flash-Lite row is the only model whose default is genuinely off.
  • thinking_budget still exists for backward compatibility, but sending thinking_budget and thinking_level in the same request "will return a 400 error". Pick one; for Gemini 3 models, pick thinking_level.

The default matters when you migrate. If you move from gemini-3-flash-preview (default high) to gemini-3.8-flash (default medium) without setting a level, your requests think less than before; if you move from 3.5 Flash to 3.1 Pro, they think more and cost more. Set the level explicitly in production so a model swap does not change behavior silently. For the older thinking_budget models and per-model migration detail, see the Gemini API Thinking Level Guide: thinking_level and thinking_budget by Model.

Setting thinkingLevel to minimal on Gemini 3 Flash with @google/genai

In the JavaScript SDK the parameter is camel-cased and lives under config.thinkingConfig. The snippet follows the shape in the @google/genai documentation and is illustrative; check the SDK reference for your installed version before using it in production:

javascript
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const response = await ai.models.generateContent({
  model: "gemini-3.5-flash", // minimal is accepted on 3.5 and 3.6 Flash, not on 3.8 Flash
  contents: "Classify this support ticket as billing, bug, or feature request: ...",
  config: {
    thinkingConfig: {
      thinkingLevel: "minimal", // "minimal" | "low" | "medium" | "high"
    },
  },
});

console.log(response.text);

The same request over REST puts the field inside generationConfig:

json
{
  "contents": [{ "parts": [{ "text": "Classify this support ticket ..." }] }],
  "generationConfig": {
    "thinkingConfig": { "thinkingLevel": "minimal" }
  }
}

minimal is not a supported value for gemini-3.8-flash; use low there, or stay on gemini-3.5-flash or gemini-3.6-flash while you still need minimal. For Pro, low is the floor. Thought summaries are off by default ("By default, only the final output is returned"); set thinking_summaries to auto if you want the model's reasoning summary alongside the answer.

Thinking tokens are billed as output: lower the level, do not cap the output

The thinking documentation is explicit that "response pricing is the sum of output tokens and thinking tokens", billed at the output rate, and that pricing "is based on the full thought tokens the model needs to generate" even when you only receive a summary or nothing at all. max_output_tokens includes thought tokens, so a tight cap can truncate the answer with an incomplete finish rather than save money. The documented way to cut cost is to lower thinking_level, not to cap output.

A worked example with stated assumptions: 1,000 requests, each 2,000 input tokens and 1,500 output tokens of which 500 are thought tokens, Standard pricing, prompts under 200k tokens.

  • gemini-3.5-flash: 2M × $1.50 + 1.5M × $9.00 = $3.00 + $13.50 = $16.50
  • gemini-3.1-pro-preview: 2M × $2.00 + 1.5M × $12.00 = $4.00 + $18.00 = $22.00
  • gemini-3-flash-preview: 2M × $0.50 + 1.5M × $3.00 = $1.00 + $4.50 = $5.50

Change the 500 thought tokens to 5,000 (a high level on a hard problem) and the output line dominates: the same 1,000 Pro requests cost $4.00 + 6M × $12.00 = $76.00. That is why the level, not the input price, usually decides your bill.

What Gemini 3 Deep Think is, who gets it, and why there is no Deep Think API

Deep Think is Google's "enhanced reasoning mode that pushes Gemini 3 performance even further", in the words of the Gemini 3 launch post of November 18, 2025. It runs on top of the Pro model in the Gemini app, and you turn it on in the prompt bar rather than in the model picker. The December 4, 2025 release note tells Ultra users to "select 'Deep Think' in the prompt bar and 'Thinking' in the model dropdown" and warns that the "response is ready – generally in a few minutes". It received a "major upgrade" on February 12, 2026.

Who gets it. Google's AI Ultra benefits page lists the "Gemini app with access to 3 Pro Deep Think (for eligible Google AI Ultra plans)" and does not state a usage count. According to the 9to5Google report of October 3, 2026, Google AI Pro "gains the Deep Think option" as part of the October 9 changes, a feature previously limited to the $99.99 and $199.99 plans; Google's model access page lists Deep Think only as a premium feature that uses more of the usage limit, without a plan column, so treat Pro access as reported rather than confirmed until you see the control in your own prompt bar.

Why you cannot call it from the API. As of October 9, 2026, no Deep Think model ID appears on the API models page, the pricing page, or the thinking documentation. "Gemini 3 Deep Think API" and "Deep Think pricing" therefore have no per-token answer: the only way to use Deep Think is a Google AI subscription, and the closest API option is gemini-3.1-pro-preview at thinking_level: "high", which is a different configuration from the mode Google benchmarked under the Deep Think name; Google has not published numbers comparing the two.

Gemini 3 Deep Think benchmarks: HLE 41.0%, GPQA Diamond 93.8%, ARC-AGI-2 45.1%

The only published Deep Think numbers are Google's own, from the November 18, 2025 launch post: Deep Think outperforms Gemini 3 Pro on "Humanity's Last Exam (41.0% without the use of tools) and GPQA Diamond (93.8%)" and scores "45.1% on ARC-AGI-2 (with code execution, ARC Prize Verified)". The February 12, 2026 upgrade shipped without a new scorecard, so those figures describe the November 2025 version against a Gemini 3 Pro that has since been replaced by 3.1 Pro in both the app and the API.

Benchmark tables that pit Gemini 3 Flash against Gemini 3 Pro (SWE-bench Verified, AIME 2025, LiveCodeBench scores from December 2025) describe a pair you can no longer choose between: gemini-3-pro-preview is shut down, and the Flash lane has moved through 3.5, 3.6, 3.7, and 3.8. For current head-to-head numbers, use Gemini 3 Flash vs Pro: Benchmarks, Pricing and Use Cases and Gemini 3.1 Pro Preview vs Gemini 3 Flash: When Pro Is Worth Paying For. Gemini 2.5 Flash, whose thinking-versus-non-thinking AIME 2025 and SWE-bench scores still circulate, remains live in the API with low, medium, and high levels; its thinking_budget controls and migration path are in the Gemini API Thinking Level Guide: thinking_level and thinking_budget by Model, and the Pro side of that generation is compared in Gemini 3.1 Pro vs Gemini 2.5 Pro: Which Model Should You Use in 2026?.

Gemini 3 Flash vs Pro in the API: 3.5 Flash and 3.8 Flash against 3.1 Pro

For developers, the Flash-versus-Pro decision comes down to three measurable differences and one you have to test yourself.

Gemini 3.5 Flash, Gemini 3 Flash preview, and Gemini 3.1 Pro API prices per million tokens

Price. Per 1M tokens on the Standard tier: gemini-3.5-flash is $1.50 input and $9.00 output; gemini-3.1-pro-preview is $2.00 and $12.00 for prompts up to 200k tokens and $4.00 and $18.00 above that; gemini-3-flash-preview is $0.50 input (text, image, video; $1.00 for audio) and $3.00 output. Context caching is $0.15 for 3.5 Flash, $0.20 or $0.40 for 3.1 Pro depending on prompt size, and $0.05 for 3 Flash preview, plus hourly storage. Prices for 3.6, 3.7, and 3.8 Flash are on the same pricing page; check them there rather than assuming they match 3.5 Flash.

Free tier. 3.5 Flash and 3 Flash preview have a free tier; 3.1 Pro does not. A prototype running on the free tier cannot simply swap its model string to Pro; you need billing enabled first. Gemini API Free Tier in 2026: What Is Free and What to Check covers the current free-tier rules.

Thinking floor. Pro cannot go below low, and neither can 3.8 or 3.7 Flash. Only 3.5 Flash, 3.6 Flash, Flash-Lite, and 3 Flash preview accept minimal. If your workload is classification, extraction, or formatting where you want the least thinking possible, that constraint alone picks the model.

Quality on your task is the part no table settles. 3.1 Pro sits at the top of the lineup and defaults to high, so it arrives already thinking hard; Flash is the model Google made the app default in December 2025 and the one it keeps iterating fastest. The practical test is to run your evaluation set on 3.8 Flash at medium and high, then on 3.1 Pro at low, and compare both quality and the thought-token count in the usage metadata. For a fuller treatment of speed and context differences, see Gemini 3 Pro vs Flash: Complete Speed and Cost Comparison Guide 2026, and for the newest Flash model specifically, Gemini 3.8 Flash API Guide: Model ID, Pricing, Limits, and Migration.

Which Gemini model to use for basic code, small data sets, or step-by-step logic

A short rule set that holds for the October 2026 lineup:

TaskPickWhy
Generating basic code, boilerplate, testsgemini-3.8-flash at medium (default)The newest Flash at its default level; raise to high only if the first pass misses edge cases
Evaluating small-to-medium data sets (CSV summaries, anomaly spotting)gemini-3.8-flash at high; move to 3.1 Pro at low if results are inconsistentFlash handles the volume; escalate to Pro only when the reasoning chain gets long
Tasks that require step-by-step logic (planning, multi-constraint scheduling, proofs)gemini-3.1-pro-preview at highThe top-end model at its default level, the deepest reasoning the API offers
Classification, extraction, routing at volumegemini-3.5-flash-lite or gemini-3.5-flash at minimalLowest thinking allowance the API offers
Questions in the app that need minutes of reasoning, not secondsDeep Think in the prompt bar (Ultra; AI Pro per reports)Nothing in the API reproduces it

Is Gemini Pro only for math and code? No. Pro is the deeper reasoning model for any domain, including long-document analysis and planning; math and code are simply where benchmark differences are easiest to measure. For writing or summarizing short texts, Flash at medium is usually indistinguishable at a fraction of the output price.

Comparing 1.5 Flash and Pro for a Python application? Those are previous generations. For a new Python service, the equivalent comparison today is gemini-3.8-flash against gemini-3.1-pro-preview, using the thinking_level table above rather than the old thinking_budget numbers. When you move to a paid tier and hit 429 errors on Pro, How to Fix Gemini API Error 429 Resource Exhausted: Complete Guide [2026] walks through the quota tiers.

Common questions about Gemini Thinking, Pro, and Deep Think

Is it better to use Gemini Thinking or Pro?

Pro, if the task is long reasoning, large documents, or code you will run, and your plan has Pro quota left. Thinking (Flash at a higher effort level in today's app) for a considered answer that comes back in seconds. Since December 17, 2025 they have been different models, not two settings of one model.

Which is better, Gemini 3.6 Thinking or 3.1 Pro?

For hard, multi-step problems, 3.1 Pro is the top-end model and defaults to the highest thinking level. For everything else, 3.6 Flash with thinking is the quicker option and leaves your Pro quota alone. Google has published no head-to-head between those two specific versions as of October 9, 2026.

Which Gemini models support thinking?

All current text models think by default. minimal, low, medium, and high are accepted by 3.5 Flash, 3.6 Flash, 3.5 Flash-Lite, and 3 Flash preview; 3.8 Flash, 3.7 Flash, and 3.1 Pro accept low, medium, and high; Gemini 2.5 models use the same three levels, with 2.5 Flash-Lite thinking off unless you turn it on.

Is there a Gemini 3 Deep Think API or API pricing?

No model ID and no per-token price are documented as of October 9, 2026. Deep Think is a Gemini app feature included in Google AI Ultra, and reportedly in AI Pro from October 2026. In the API, gemini-3.1-pro-preview at thinking_level: "high" is the deepest reasoning you can request.

Is Gemini 3 Pro still available in the API?

No. gemini-3-pro-preview is marked shut down on the models page; use gemini-3.1-pro-preview, which accepts low, medium, and high and costs $2.00/$12.00 per 1M tokens for prompts up to 200k tokens.

Why does my Gemini 3 request return a 400 error when I set thinking?

Check whether thinking_budget and thinking_level are both present. The developer guide says using them together "will return a 400 error". Remove thinking_budget on Gemini 3 models, and confirm the level you send is in that model's supported list; minimal is not supported on 3.8 Flash or 3.1 Pro.