# Gemini 3 Deep Think vs Flash vs Pro: Thinking Levels Explained

> Deep Think is an app-only mode with no API model; Flash and Pro are app options and API models with a per-request thinking_level. Levels, prices, plan access.

- Source: https://www.aifreeapi.com/en/posts/gemini-3-deep-think-vs-flash-vs-pro
- Language: en
- Published: 2026-01-04
- Updated: 2026-10-09
- Publisher: AI Free API (https://www.aifreeapi.com)

"Deep Think", "Flash", and "Pro" are not three models on one menu. Deep Think is a reasoning mode inside the Gemini app, available to Google AI Ultra subscribers (and, according to reports, to AI Pro subscribers from October 2026), and it takes minutes per answer. Flash and Pro exist in two places at once: as options in the app's model picker and as API models (`gemini-3.8-flash`, `gemini-3.5-flash`, `gemini-3.1-pro-preview`) whose reasoning depth you set per request with `thinking_level`. There is no Deep Think model ID in the API models list, the pricing page, or the thinking documentation as of October 9, 2026.

If you are in the app and torn between Thinking and Pro, the short answer is: Pro is the Gemini 3.1 Pro model, and the effort level you pick decides how long it thinks. If you are in the API and want Flash to stop thinking as much as possible, set `thinking_level` to `minimal` on `gemini-3.5-flash` or `gemini-3.6-flash`; the newest `gemini-3.8-flash` only goes down to `low`. The sections below give the full per-model table, the plan changes that start today, and what Deep Think actually gets you.

## Deep Think vs Flash vs Pro at a glance: one table across app and API

Everything else in this guide expands on this one table. Prices are Standard tier, per 1M tokens, from the [Gemini API pricing page](https://ai.google.dev/gemini-api/docs/pricing); levels and defaults are from the [thinking documentation](https://ai.google.dev/gemini-api/docs/thinking); plan access is from Google's [model access help page](https://support.google.com/gemini/answer/17004136), all as of October 9, 2026.

| Name you see | What it is | API model ID | `thinking_level` values (default) | API price in / out | Who can use it in the app |
| --- | --- | --- | --- | --- | --- |
| Flash-Lite | Smallest Flash | `gemini-3.5-flash-lite` | minimal / low / medium / high (minimal) | listed on the pricing page | Everyone, including no subscription |
| Flash (shown as "Fast" and "Thinking" in late 2025) | The default workhorse model | `gemini-3.8-flash` (newest stable), `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash` (legacy) | 3.8 and 3.7: low / medium / high (medium); 3.6 and 3.5: minimal / low / medium / high (medium) | 3.5 Flash: $1.50 / $9.00, free tier available | AI Plus, AI Pro, AI Ultra |
| Pro | The top-end model | `gemini-3.1-pro-preview` | low / medium / high (high); no `minimal` | $2.00 / $12.00 up to 200k prompt tokens; $4.00 / $18.00 above; no free tier | AI Pro, AI Ultra |
| Deep Think | A reasoning mode on top of Pro, answers in minutes | none documented | not applicable | no API price; included in the subscription | AI Ultra; AI Pro according to October 2026 reports |

![Gemini app picker labels mapped to Gemini API model IDs and thinking levels](https://www.aifreeapi.com/posts/en/gemini-3-deep-think-vs-flash-vs-pro/img/comparison.png)

Two things in this table surprise most people. First, "Gemini 3 Pro" is no longer an API model: `gemini-3-pro-preview` is listed as shut down on the [models page](https://ai.google.dev/gemini-api/docs/models), and the Pro lane is `gemini-3.1-pro-preview`. Second, the cheapest "Flash" price you may remember, $0.50 input and $3.00 output, belongs to `gemini-3-flash-preview`, which is still available but is a preview model from December 2025; the current 3.5 Flash costs $1.50 and $9.00.

## Gemini Thinking vs Pro in the app: what Fast, Thinking, and Pro mean

In the Gemini app, "Thinking" has meant two different things, which is why the question keeps coming up. According to the [Gemini app release notes](https://gemini.google/release-notes/):

- **November 18, 2025**: Gemini 3 Pro arrived for everyone in the app, and you reached it by "selecting 'Thinking' in the model drop down". For a month, Thinking *was* Pro.
- **December 17, 2025**: Gemini 3 Flash became the default. The picker changed to "Fast" for quick answers and "Thinking" for harder problems, both running on Gemini 3 Flash, while Gemini 3 Pro moved to its own "Pro" entry. From this point on, Thinking was Flash with more reasoning, and Pro was the bigger model.
- **February 19, 2026**: Gemini 3.1 Pro rolled out behind the "Pro" entry, with higher limits for AI Pro and AI Ultra.
- **May 19, 2026 and July 21, 2026**: 3.5 Flash and then 3.6 Flash became selectable in the model drop-down.

So "Flash vs Thinking vs Pro" from the December 2025 picker maps to: Fast = Gemini 3 Flash answering quickly, Thinking = Gemini 3 Flash with extended reasoning, Pro = Gemini 3 Pro (later 3.1 Pro). Any "Thinking vs Pro" comparison from that period is comparing Flash with extended reasoning against the Pro model, not two tiers of one model.

The current app no longer uses the Fast/Thinking split as the main choice. Google's [model access page](https://support.google.com/gemini/answer/17004136) describes a model list of Flash-Lite, Flash, and Pro, and for each available model you can select an effort level of low, medium, or high. That effort level is the consumer-side counterpart of the API's `thinking_level`. "Extended thinking and Deep Think" are listed as premium features that use more of your usage limit, so a Pro answer at high effort uses up your plan's limit faster than a Flash answer at low effort.

**Thinking or Pro for your question?** Pick Pro when the task is long multi-step reasoning, large documents, or code you will actually run; pick Flash (the former Thinking) at medium or high effort when you want a considered answer quickly and your plan's Pro quota is the scarce resource. For everyday questions, Flash at low effort or Flash-Lite is the right default, and that is also the only thing a non-subscriber gets from today.

### Which Gemini models each plan gets from October 9, 2026

Google's model access page says the changes "will start to take effect for users without an AI subscription on October 9th", and that AI Plus subscribers will get an email with their date. The table on that page, plus the AI Ultra benefits page and a [9to5Google report of October 3, 2026](https://9to5google.com/2026/10/04/gemini-model-limits-oct-26/) for the parts Google has not written down, gives this picture:

| Plan | Models in the Gemini app | Deep Think |
| --- | --- | --- |
| No subscription | Flash-Lite only (Flash and Pro removed from October 9, 2026) | No |
| Google AI Plus ($4.99/month per the report) | Flash-Lite and Flash; Pro removed on a date sent by email | No |
| Google AI Pro ($19.99/month per the report) | Flash-Lite, Flash, Pro | According to the report, gains the Deep Think option; Google's page does not list it |
| Google AI Ultra (starting at $100/month per the May 19, 2026 release note) | Flash-Lite, Flash, Pro, with 5x or 20x the usage quota of AI Pro | Yes, "3 Pro Deep Think" for eligible Ultra plans |

Google's own page does not say which Flash version each plan gets; the report names 3.6 Flash and 3.5 Flash-Lite as the models involved. If you are on AI Plus and rely on Pro, treat the email date as your deadline rather than today.

## Gemini 3 thinking levels in the API: minimal, low, medium, high by model

Every current Gemini text model thinks by default; `thinking_level` controls how much. The values are `minimal`, `low`, `medium`, and `high`, but not every model accepts every value, and the default differs by model. This table is the current one from the [thinking documentation](https://ai.google.dev/gemini-api/docs/thinking) as of October 9, 2026:

| Model | Supported `thinking_level` | Default |
| --- | --- | --- |
| `gemini-3.8-flash`, `gemini-3.7-flash` | low, medium, high | medium |
| `gemini-3.6-flash`, `gemini-3.5-flash` | minimal, low, medium, high | medium |
| `gemini-3.5-flash-lite` | minimal, low, medium, high | minimal |
| `gemini-3.1-pro-preview` | low, medium, high | high |
| `gemini-3-flash-preview` | minimal, low, medium, high | high |
| `gemini-2.5-pro`, `gemini-2.5-flash` | low, medium, high | thinking on |
| `gemini-2.5-flash-lite` | low, medium, high | thinking off |

Three rules from the [Gemini 3 developer guide](https://ai.google.dev/gemini-api/docs/gemini-3) change how you should read that table:

- Levels are "relative allowances for thinking rather than strict token guarantees". `high` on Flash-Lite is not the same number of thought tokens as `high` on Pro.
- `minimal` "does not guarantee that thinking is off". It is the lowest allowance, not a switch. The 2.5 Flash-Lite row is the only model whose default is genuinely off.
- `thinking_budget` still exists for backward compatibility, but sending `thinking_budget` and `thinking_level` in the same request "will return a 400 error". Pick one; for Gemini 3 models, pick `thinking_level`.

The default matters when you migrate. If you move from `gemini-3-flash-preview` (default `high`) to `gemini-3.8-flash` (default `medium`) without setting a level, your requests think less than before; if you move from 3.5 Flash to 3.1 Pro, they think more and cost more. Set the level explicitly in production so a model swap does not change behavior silently. For the older `thinking_budget` models and per-model migration detail, see the [Gemini API Thinking Level Guide: thinking_level and thinking_budget by Model](/en/posts/gemini-api-thinking-level).

### Setting thinkingLevel to minimal on Gemini 3 Flash with @google/genai

In the JavaScript SDK the parameter is camel-cased and lives under `config.thinkingConfig`. The snippet follows the shape in the `@google/genai` documentation and is illustrative; check the SDK reference for your installed version before using it in production:

```javascript
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const response = await ai.models.generateContent({
  model: "gemini-3.5-flash", // minimal is accepted on 3.5 and 3.6 Flash, not on 3.8 Flash
  contents: "Classify this support ticket as billing, bug, or feature request: ...",
  config: {
    thinkingConfig: {
      thinkingLevel: "minimal", // "minimal" | "low" | "medium" | "high"
    },
  },
});

console.log(response.text);
```

The same request over REST puts the field inside `generationConfig`:

```json
{
  "contents": [{ "parts": [{ "text": "Classify this support ticket ..." }] }],
  "generationConfig": {
    "thinkingConfig": { "thinkingLevel": "minimal" }
  }
}
```

`minimal` is not a supported value for `gemini-3.8-flash`; use `low` there, or stay on `gemini-3.5-flash` or `gemini-3.6-flash` while you still need `minimal`. For Pro, `low` is the floor. Thought summaries are off by default ("By default, only the final output is returned"); set `thinking_summaries` to `auto` if you want the model's reasoning summary alongside the answer.

### Thinking tokens are billed as output: lower the level, do not cap the output

The thinking documentation is explicit that "response pricing is the sum of output tokens and thinking tokens", billed at the output rate, and that pricing "is based on the full thought tokens the model needs to generate" even when you only receive a summary or nothing at all. `max_output_tokens` includes thought tokens, so a tight cap can truncate the answer with an `incomplete` finish rather than save money. The documented way to cut cost is to lower `thinking_level`, not to cap output.

A worked example with stated assumptions: 1,000 requests, each 2,000 input tokens and 1,500 output tokens of which 500 are thought tokens, Standard pricing, prompts under 200k tokens.

- `gemini-3.5-flash`: 2M × $1.50 + 1.5M × $9.00 = $3.00 + $13.50 = **$16.50**
- `gemini-3.1-pro-preview`: 2M × $2.00 + 1.5M × $12.00 = $4.00 + $18.00 = **$22.00**
- `gemini-3-flash-preview`: 2M × $0.50 + 1.5M × $3.00 = $1.00 + $4.50 = **$5.50**

Change the 500 thought tokens to 5,000 (a `high` level on a hard problem) and the output line dominates: the same 1,000 Pro requests cost $4.00 + 6M × $12.00 = $76.00. That is why the level, not the input price, usually decides your bill.

## What Gemini 3 Deep Think is, who gets it, and why there is no Deep Think API

Deep Think is Google's "enhanced reasoning mode that pushes Gemini 3 performance even further", in the words of the [Gemini 3 launch post](https://blog.google/products-and-platforms/products/gemini/gemini-3/) of November 18, 2025. It runs on top of the Pro model in the Gemini app, and you turn it on in the prompt bar rather than in the model picker. The December 4, 2025 release note tells Ultra users to "select 'Deep Think' in the prompt bar and 'Thinking' in the model dropdown" and warns that the "response is ready – generally in a few minutes". It received a "major upgrade" on February 12, 2026.

**Who gets it.** Google's [AI Ultra benefits page](https://support.google.com/googleone/answer/16286513) lists the "Gemini app with access to 3 Pro Deep Think (for eligible Google AI Ultra plans)" and does not state a usage count. According to the 9to5Google report of October 3, 2026, Google AI Pro "gains the Deep Think option" as part of the October 9 changes, a feature previously limited to the $99.99 and $199.99 plans; Google's model access page lists Deep Think only as a premium feature that uses more of the usage limit, without a plan column, so treat Pro access as reported rather than confirmed until you see the control in your own prompt bar.

**Why you cannot call it from the API.** As of October 9, 2026, no Deep Think model ID appears on the API models page, the pricing page, or the thinking documentation. "Gemini 3 Deep Think API" and "Deep Think pricing" therefore have no per-token answer: the only way to use Deep Think is a Google AI subscription, and the closest API option is `gemini-3.1-pro-preview` at `thinking_level: "high"`, which is a different configuration from the mode Google benchmarked under the Deep Think name; Google has not published numbers comparing the two.

### Gemini 3 Deep Think benchmarks: HLE 41.0%, GPQA Diamond 93.8%, ARC-AGI-2 45.1%

The only published Deep Think numbers are Google's own, from the November 18, 2025 launch post: Deep Think outperforms Gemini 3 Pro on "Humanity's Last Exam (41.0% without the use of tools) and GPQA Diamond (93.8%)" and scores "45.1% on ARC-AGI-2 (with code execution, ARC Prize Verified)". The February 12, 2026 upgrade shipped without a new scorecard, so those figures describe the November 2025 version against a Gemini 3 Pro that has since been replaced by 3.1 Pro in both the app and the API.

Benchmark tables that pit Gemini 3 Flash against Gemini 3 Pro (SWE-bench Verified, AIME 2025, LiveCodeBench scores from December 2025) describe a pair you can no longer choose between: `gemini-3-pro-preview` is shut down, and the Flash lane has moved through 3.5, 3.6, 3.7, and 3.8. For current head-to-head numbers, use [Gemini 3 Flash vs Pro: Benchmarks, Pricing and Use Cases](/en/posts/gemini-3-flash-vs-pro-capabilities) and [Gemini 3.1 Pro Preview vs Gemini 3 Flash: When Pro Is Worth Paying For](/en/posts/gemini-3-1-pro-preview-vs-gemini-3-flash). Gemini 2.5 Flash, whose thinking-versus-non-thinking AIME 2025 and SWE-bench scores still circulate, remains live in the API with `low`, `medium`, and `high` levels; its `thinking_budget` controls and migration path are in the [Gemini API Thinking Level Guide: thinking_level and thinking_budget by Model](/en/posts/gemini-api-thinking-level), and the Pro side of that generation is compared in [Gemini 3.1 Pro vs Gemini 2.5 Pro: Which Model Should You Use in 2026?](/en/posts/gemini-3-1-pro-vs-gemini-2-5-pro).

## Gemini 3 Flash vs Pro in the API: 3.5 Flash and 3.8 Flash against 3.1 Pro

For developers, the Flash-versus-Pro decision comes down to three measurable differences and one you have to test yourself.

![Gemini 3.5 Flash, Gemini 3 Flash preview, and Gemini 3.1 Pro API prices per million tokens](https://www.aifreeapi.com/posts/en/gemini-3-deep-think-vs-flash-vs-pro/img/pricing.png)

**Price.** Per 1M tokens on the Standard tier: `gemini-3.5-flash` is $1.50 input and $9.00 output; `gemini-3.1-pro-preview` is $2.00 and $12.00 for prompts up to 200k tokens and $4.00 and $18.00 above that; `gemini-3-flash-preview` is $0.50 input (text, image, video; $1.00 for audio) and $3.00 output. Context caching is $0.15 for 3.5 Flash, $0.20 or $0.40 for 3.1 Pro depending on prompt size, and $0.05 for 3 Flash preview, plus hourly storage. Prices for 3.6, 3.7, and 3.8 Flash are on the same pricing page; check them there rather than assuming they match 3.5 Flash.

**Free tier.** 3.5 Flash and 3 Flash preview have a free tier; 3.1 Pro does not. A prototype running on the free tier cannot simply swap its model string to Pro; you need billing enabled first. [Gemini API Free Tier in 2026: What Is Free and What to Check](/en/posts/google-gemini-api-free-tier) covers the current free-tier rules.

**Thinking floor.** Pro cannot go below `low`, and neither can 3.8 or 3.7 Flash. Only 3.5 Flash, 3.6 Flash, Flash-Lite, and 3 Flash preview accept `minimal`. If your workload is classification, extraction, or formatting where you want the least thinking possible, that constraint alone picks the model.

**Quality on your task** is the part no table settles. 3.1 Pro sits at the top of the lineup and defaults to `high`, so it arrives already thinking hard; Flash is the model Google made the app default in December 2025 and the one it keeps iterating fastest. The practical test is to run your evaluation set on 3.8 Flash at `medium` and `high`, then on 3.1 Pro at `low`, and compare both quality and the thought-token count in the usage metadata. For a fuller treatment of speed and context differences, see [Gemini 3 Pro vs Flash: Complete Speed and Cost Comparison Guide 2026](/en/posts/gemini-3-pro-vs-flash-speed-cost), and for the newest Flash model specifically, [Gemini 3.8 Flash API Guide: Model ID, Pricing, Limits, and Migration](/en/posts/gemini-3-8-flash).

## Which Gemini model to use for basic code, small data sets, or step-by-step logic

A short rule set that holds for the October 2026 lineup:

| Task | Pick | Why |
| --- | --- | --- |
| Generating basic code, boilerplate, tests | `gemini-3.8-flash` at `medium` (default) | The newest Flash at its default level; raise to `high` only if the first pass misses edge cases |
| Evaluating small-to-medium data sets (CSV summaries, anomaly spotting) | `gemini-3.8-flash` at `high`; move to 3.1 Pro at `low` if results are inconsistent | Flash handles the volume; escalate to Pro only when the reasoning chain gets long |
| Tasks that require step-by-step logic (planning, multi-constraint scheduling, proofs) | `gemini-3.1-pro-preview` at `high` | The top-end model at its default level, the deepest reasoning the API offers |
| Classification, extraction, routing at volume | `gemini-3.5-flash-lite` or `gemini-3.5-flash` at `minimal` | Lowest thinking allowance the API offers |
| Questions in the app that need minutes of reasoning, not seconds | Deep Think in the prompt bar (Ultra; AI Pro per reports) | Nothing in the API reproduces it |

**Is Gemini Pro only for math and code?** No. Pro is the deeper reasoning model for any domain, including long-document analysis and planning; math and code are simply where benchmark differences are easiest to measure. For writing or summarizing short texts, Flash at `medium` is usually indistinguishable at a fraction of the output price.

**Comparing 1.5 Flash and Pro for a Python application?** Those are previous generations. For a new Python service, the equivalent comparison today is `gemini-3.8-flash` against `gemini-3.1-pro-preview`, using the `thinking_level` table above rather than the old `thinking_budget` numbers. When you move to a paid tier and hit 429 errors on Pro, [How to Fix Gemini API Error 429 Resource Exhausted: Complete Guide [2026]](/en/posts/gemini-api-error-429-resource-exhausted-fix) walks through the quota tiers.

## Common questions about Gemini Thinking, Pro, and Deep Think

### Is it better to use Gemini Thinking or Pro?

Pro, if the task is long reasoning, large documents, or code you will run, and your plan has Pro quota left. Thinking (Flash at a higher effort level in today's app) for a considered answer that comes back in seconds. Since December 17, 2025 they have been different models, not two settings of one model.

### Which is better, Gemini 3.6 Thinking or 3.1 Pro?

For hard, multi-step problems, 3.1 Pro is the top-end model and defaults to the highest thinking level. For everything else, 3.6 Flash with thinking is the quicker option and leaves your Pro quota alone. Google has published no head-to-head between those two specific versions as of October 9, 2026.

### Which Gemini models support thinking?

All current text models think by default. `minimal`, `low`, `medium`, and `high` are accepted by 3.5 Flash, 3.6 Flash, 3.5 Flash-Lite, and 3 Flash preview; 3.8 Flash, 3.7 Flash, and 3.1 Pro accept `low`, `medium`, and `high`; Gemini 2.5 models use the same three levels, with 2.5 Flash-Lite thinking off unless you turn it on.

### Is there a Gemini 3 Deep Think API or API pricing?

No model ID and no per-token price are documented as of October 9, 2026. Deep Think is a Gemini app feature included in Google AI Ultra, and reportedly in AI Pro from October 2026. In the API, `gemini-3.1-pro-preview` at `thinking_level: "high"` is the deepest reasoning you can request.

### Is Gemini 3 Pro still available in the API?

No. `gemini-3-pro-preview` is marked shut down on the models page; use `gemini-3.1-pro-preview`, which accepts `low`, `medium`, and `high` and costs $2.00/$12.00 per 1M tokens for prompts up to 200k tokens.

### Why does my Gemini 3 request return a 400 error when I set thinking?

Check whether `thinking_budget` and `thinking_level` are both present. The developer guide says using them together "will return a 400 error". Remove `thinking_budget` on Gemini 3 models, and confirm the level you send is in that model's supported list; `minimal` is not supported on 3.8 Flash or 3.1 Pro.
