# Chat Completions vs Responses API: Which One Should You Use?

> Chat Completions is still supported, but Responses is OpenAI's pick for new work and the only way to pair tool calls with reasoning on GPT-5.4 and later.

- Source: https://www.aifreeapi.com/en/posts/chat-completions-vs-responses-api
- Language: en
- Published: 2026-08-25
- Updated: 2026-09-25
- Publisher: AI Free API (https://www.aifreeapi.com)

## Use Responses for new work; keep Chat Completions where it works

Use the Responses API (`/v1/responses`) for anything new, and for any flow where a reasoning model has to call tools. Keep Chat Completions (`/v1/chat/completions`) where it already works for plain text or Structured Outputs, where you need several candidates per request, or where the same code must also run against providers that only speak Chat Completions.

Chat Completions is not deprecated. OpenAI's [migration guide](https://developers.openai.com/api/docs/guides/migrate-to-responses) says it "remains supported" while Responses "is recommended for all new projects," and that you can move one user flow at a time. What changed in 2026 is the model side: starting with GPT-5.4, Chat Completions only allows tool calling when `reasoning_effort` is `none`, and GPT-6 Astra can't call functions through Chat Completions at all.

| Your situation | Use | Deciding reason |
|---|---|---|
| Function calling with reasoning on GPT-5.4 or later | Responses | Chat Completions allows tool calls only with `reasoning_effort: "none"` |
| Function calling on GPT-6 Astra | Responses | Astra requires Responses for tool calling |
| OpenAI-hosted tools (web search, file search, code interpreter, remote MCP, image generation, computer use) | Responses | Chat Completions has no native hosted tools |
| Server-side conversation state | Responses | `previous_response_id` or the Conversations API |
| One code path for OpenAI and "OpenAI-compatible" providers | Chat Completions, or both | Many providers expose Chat Completions; Responses support varies |
| Several candidate answers per request (`n` > 1) | Chat Completions | Responses removed `n` |
| Any other new project on OpenAI models | Responses | OpenAI's recommendation for all new projects |
| A text-only or Structured Outputs flow that already works | Stay, migrate later | Supported, with no shutdown date |

Read the rows from the top and stop at the first one that fits your flow. The decision is per flow, not per app: if one flow needs reasoning plus tools, that flow moves even if the rest stays put.

![Decision flow with four questions in order: tools with reasoning on GPT-5.4 or later or GPT-6 Astra leads to Responses; hosted tools or server-side state leads to Responses; a Chat-only provider or n greater than 1 leads to Chat Completions; a new flow leads to Responses; no to all four means keep Chat Completions and migrate later](https://www.aifreeapi.com/posts/en/chat-completions-vs-responses-api/img/which-api-decision-flow.webp)

## Is Chat Completions deprecated?

No. As of September 25, 2026, OpenAI's [deprecations page](https://developers.openai.com/api/docs/deprecations) has no entry that retires `/v1/chat/completions`. The confusion comes from three different APIs with similar names:

| API | Endpoint | Status as of September 25, 2026 |
|---|---|---|
| Completions (legacy) | `/v1/completions` | The older prompt-in, text-out API. Model pages label it "Completions (legacy)." |
| Assistants API | `/v1/assistants` | Shut down on August 26, 2026. OpenAI names the Responses API and the Conversations API as its replacement. |
| Chat Completions | `/v1/chat/completions` | Supported. No deprecation entry. |
| Responses | `/v1/responses` | Recommended for all new projects. |

OpenAI gave Assistants users notice on August 26, 2025, and removed the API exactly one year later. If your code still calls `/v1/assistants`, it needs the Responses and Conversations APIs now, not later.

One more source of doubt: OpenAI's own Codex CLI dropped its `chat/completions` wire support in February 2026, and [the announcement](https://github.com/openai/codex/discussions/7782) describes that protocol as one that "originated in the GPT-3.5 era." That decision covers Codex as a client. Custom providers in Codex had to switch from `wire_api = "chat"` to `wire_api = "responses"`, but the API endpoint itself kept running. If that is what brought you here, the practical fix is in [Connect Codex to a Third-Party Model API Without Guessing the Protocol](/en/posts/codex-third-party-api).

## The 2026 rule that decides most cases: tools plus reasoning

If your app calls tools, a model constraint can settle the question before any difference in API design does. OpenAI states it in three guides:

- The migration guide: "Starting with GPT-5.4, Chat Completions does not support tool calling with reasoning_effort values other than none."
- The [reasoning guide](https://developers.openai.com/api/docs/guides/reasoning) and the [function calling guide](https://developers.openai.com/api/docs/guides/function-calling): "GPT-6 Astra requires the Responses API for tool calling." The reasoning guide adds that Astra rejects `none` reasoning effort with HTTP 400.

Put together, this is how tool calling looks per API for current models:

| Model | Chat Completions + tools | Responses + tools |
|---|---|---|
| GPT-6 Astra (`gpt-6-astra`) | Not supported. Astra accepts `low` through `max` effort, never `none` | Supported, including hosted tools listed on the [model page](https://developers.openai.com/api/docs/models/gpt-6-astra) |
| GPT-5.4 and later, for example GPT-5.6 Sol (`gpt-5.6`, alias of `gpt-5.6-sol`) or GPT-6 Sol (`gpt-6-sol`) | Only with `reasoning_effort: "none"` | Supported at any effort the model allows |
| Models before GPT-5.4 | Not restricted by this rule | Supported |

The middle row applies OpenAI's "starting with GPT-5.4" wording to the newer models; it is a reading of the rule, not a separate test. The [GPT-5.6 Sol page](https://developers.openai.com/api/docs/models/gpt-5.6-sol) lists `medium` as its default effort, so a Chat Completions request that sends tools to `gpt-5.6` has to turn reasoning off. On Responses, the same request keeps reasoning and tools together. If you aren't sure which model family your app uses, [OpenAI Text Models in 2026: GPT-5.6 Sol, Terra, and Luna Compared](/en/posts/openai-text-models-2026) lists the current IDs.

OpenAI also claims quality and cost gains from Responses with reasoning models. From its internal evals: a 3% improvement on SWE-bench "with same prompt and setup," and 40% to 80% better cache utilization compared with Chat Completions. Those are vendor numbers, not an independent benchmark, and cache utilization doesn't translate into a guaranteed saving on your bill.

## What changes in your code when you move to Responses

Most of the migration is a handful of renamed fields plus one real design change: output and input become arrays of typed **Items** (`message`, `reasoning`, `function_call`, `function_call_output`, and so on) instead of chat messages that bundle everything together.

| Concern | Chat Completions | Responses |
|---|---|---|
| Endpoint | `POST /v1/chat/completions` | `POST /v1/responses` |
| SDK call | `client.chat.completions.create` | `client.responses.create` |
| Input | `messages` | `input` (a string or an array of Items) |
| System or developer guidance | A `system` or `developer` message | Top-level `instructions`, or a compatible message Item |
| Where the text lives | `choices[0].message.content` | `output_text` (SDK helper) or the Items in `output` |
| Multiple candidates | `n` | Not available; make separate requests |
| Function definition | Externally tagged: `{"type": "function", "function": {...}}` | Internally tagged: `{"type": "function", "name": ..., ...}` |
| Function strictness | Non-strict unless you set `strict` | Tries strict by default, falls back to non-strict |
| Tool call and result | `tool_calls` on the assistant message, then a `role: "tool"` message with `tool_call_id` | `function_call` Item, then a `function_call_output` Item with the same `call_id` |
| Structured Outputs | `response_format` | `text.format` |
| Streaming | Chunks with a `delta` field | Typed server-sent events |
| Multi-turn state | You resend the whole `messages` array | `previous_response_id`, manual Item replay, or the Conversations API |

### Basic text generation

A simple role/content array works as Responses `input` without changes, as long as it doesn't contain function calls or multimodal parts. The part that breaks is reading the result.

```python
from openai import OpenAI

client = OpenAI()

messages = [
    {"role": "system", "content": "Answer in one sentence."},
    {"role": "user", "content": "What is a webhook?"},
]

# Chat Completions
completion = client.chat.completions.create(model="gpt-5.6", messages=messages)
print(completion.choices[0].message.content)

# Responses: same array as input
response = client.responses.create(model="gpt-5.6", input=messages)
print(response.output_text)

# Responses: guidance moved to instructions
response = client.responses.create(
    model="gpt-5.6",
    instructions="Answer in one sentence.",
    input="What is a webhook?",
)
print(response.output_text)
```

`output_text` joins the final text for you. When a response can contain reasoning or tool calls, loop over `response.output` and branch on each Item's `type`; the first Item is not always a message.

### Function calling

![Tool-call round trip in two lanes: Chat Completions sends messages with tools defined under function, gets message.tool_calls each with an id, and replies with a role tool message carrying tool_call_id; Responses sends input with tools named at the top level, gets reasoning and function_call Items with call_id, and sends back response.output plus function_call_output](https://www.aifreeapi.com/posts/en/chat-completions-vs-responses-api/img/tool-call-round-trip.webp)

The tool definition loses its nested `function` object. Also check strictness: Chat Completions functions are non-strict by default, while Responses attempts strict mode when you omit `strict`. If the schema can't be made strict, Responses falls back to best-effort calling and reports `strict: false` on the resolved tool. Set `"strict": False` yourself if you want the old behavior without surprises.

```python
import json

# Chat Completions definition (externally tagged)
chat_tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    },
}]

# Responses definition (internally tagged)
tools = [{
    "type": "function",
    "name": "get_weather",
    "description": "Get current weather for a city.",
    "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
        "additionalProperties": False,
    },
}]

def get_weather(city: str) -> dict:
    return {"city": city, "forecast": "light rain"}  # replace with a real lookup

input_items = [{"role": "user", "content": "Do I need an umbrella in Paris today?"}]
response = client.responses.create(model="gpt-5.6", input=input_items, tools=tools)

# Keep every returned Item, including reasoning, before adding results
input_items += response.output
for item in response.output:
    if item.type == "function_call":
        args = json.loads(item.arguments)
        input_items.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(get_weather(**args)),
        })

final = client.responses.create(model="gpt-5.6", input=input_items, tools=tools)
print(final.output_text)
```

Two details matter with reasoning models. First, the function calling guide says reasoning Items returned alongside tool calls "must also be passed back with tool call outputs," which is why the loop appends all of `response.output` and not just the calls. Second, the model can return several calls in one turn. Handle each one, or set `parallel_tool_calls` to `false` to get zero or one call.

### Structured Outputs

Both APIs support strict JSON Schema output. Only the location of the schema moves, from `response_format` to `text.format`, and the `json_schema` wrapper is flattened.

```python
schema = {
    "type": "object",
    "properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
    "required": ["name", "age"],
    "additionalProperties": False,
}

# Chat Completions
client.chat.completions.create(
    model="gpt-5.6",
    messages=[{"role": "user", "content": "Jane, 54, joined in March."}],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "person", "strict": True, "schema": schema},
    },
)

# Responses
client.responses.create(
    model="gpt-5.6",
    input="Jane, 54, joined in March.",
    text={"format": {"type": "json_schema", "name": "person", "strict": True, "schema": schema}},
)
```

Sending `response_format` to `/v1/responses` is one of the mistakes OpenAI lists in its migration guide.

### Streaming

A Chat Completions stream is a series of chunks with `choices[0].delta`. A Responses stream is a series of typed events, so the handler has to branch on `event.type`.

```python
# Chat Completions
stream = client.chat.completions.create(model="gpt-5.6", messages=messages, stream=True)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

# Responses
stream = client.responses.create(model="gpt-5.6", input=messages, stream=True)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="")
    elif event.type == "response.completed":
        usage = event.response.usage
    elif event.type == "error":
        raise RuntimeError(event.message)
```

The events OpenAI names for text streaming are `response.created`, `response.output_text.delta`, `response.completed`, and `error`. Function calls add `response.function_call_arguments.delta` and `response.function_call_arguments.done`. The [streaming guide](https://developers.openai.com/api/docs/guides/streaming-responses) covers the other event types, such as `response.failed` and `response.output_item.added`.

## Conversation state, storage, and billing

With Chat Completions, your application stores the transcript and sends the accumulated `messages` array every turn. The Responses API gives you [three options](https://developers.openai.com/api/docs/guides/conversation-state):

1. `previous_response_id`: point the next request at the last response and OpenAI supplies the prior context.
2. Manual Item replay: pass earlier output Items back in `input`, which lets you trim or edit context yourself.
3. The Conversations API: a persistent conversation object that lives until you delete it.

```python
first = client.responses.create(
    model="gpt-5.6",
    instructions="You are a concise support agent.",
    input="My invoice shows two charges for September.",
)

follow_up = client.responses.create(
    model="gpt-5.6",
    previous_response_id=first.id,
    instructions="You are a concise support agent.",  # not carried over, so resend it
    input="Can you refund the duplicate?",
)
```

Keep two limits in mind with `previous_response_id`. It doesn't carry over the previous response's top-level `instructions`, so resend them on every request. It also doesn't make earlier turns free: OpenAI states that "all previous input tokens for responses in the chain are billed as input tokens." Any cost benefit comes from caching, a separate mechanism covered in [OpenAI vs Claude Prompt Caching Cost: Cached Token Pricing and Break-Even Math](/en/topics/claude).

Storage defaults differ as well. According to OpenAI's [data controls page](https://developers.openai.com/api/docs/guides/your-data):

- Responses are stored by default, with a 30-day application state retention period by default or when `store` is `true`. Set `store: false` to turn that off.
- Conversation objects are not subject to the 30-day limit; they stay until deleted, and `/v1/conversations` is not eligible for Zero Data Retention (ZDR).
- For ZDR organizations, `store` is always treated as `false`. To keep reasoning context across turns without stored state, replay the encrypted reasoning Items the API returns (each carries `encrypted_content`).
- For `/v1/chat/completions`, the data controls table lists no application state retention except audio outputs (1 hour). The migration guide, however, says "Chat completions are stored by default for new accounts."

Those two Chat Completions statements describe different kinds of storage. If you don't want outputs kept, set `store: false` explicitly on both APIs instead of relying on a default.

## When staying on Chat Completions is the right call

Moving has a cost, and in several situations it buys you nothing yet:

- **The flow is plain text or Structured Outputs, with no tools.** It works, it's supported, and there's no shutdown date. Migrate it when you next touch that code.
- **You need `n` > 1.** Responses returns one generation per request. Getting three candidates means three separate requests and your own fan-out logic.
- **Your code has to stay portable.** Many third-party "OpenAI-compatible" endpoints implement Chat Completions; Responses support ranges from none to partial. Some providers accept Responses-style requests but ignore unsupported parameters and tools and keep no state. Check a named provider's docs before you assume parity.
- **You use tools on GPT-5.4 or later and don't need reasoning for them.** A simple extraction or routing call can run with `reasoning_effort: "none"` on Chat Completions. The moment that flow needs reasoning, it has to move.

Staying doesn't mean freezing. OpenAI recommends "migrating all flows to the Responses API over time to take advantage of the latest OpenAI features and improvements," so keep the migration on your roadmap even for flows that can wait.

## How to migrate one flow and prove it works

OpenAI's rollout checklist works best when you treat each step as a test that has to pass before more traffic moves. A workable order:

1. **Start with the simplest text flow.** Change the endpoint and request body, then replace every read of `choices[0].message.content` with `output_text` or an Item loop.
2. **Pick one state strategy for that flow**: `previous_response_id`, manual Item replay, or Conversations. Resend `instructions` if you chain responses.
3. **Set `store` deliberately.** For stateless or ZDR flows, use `store: false` and replay encrypted reasoning Items.
4. **Port function definitions and the tool loop.** Verify that every `function_call_output` carries the matching `call_id` and that reasoning Items travel with the results. Run a schema that used to be non-strict to see whether Responses accepts it strictly or falls back.
5. **Move Structured Outputs** from `response_format` to `text.format`.
6. **Rewrite the stream handler** for typed events, including `error` and `response.failed`.
7. **Replace custom tool plumbing with hosted tools** where they fit, such as web search or file search.
8. **Compare before shifting traffic.** OpenAI's checklist ends with comparing behavior, latency, token usage, and errors between the two paths. Usage fields are named differently on each API; [Reasoning Token Billing: Avoid Double Counting Across Gemini, OpenAI, and Claude](/en/posts/reasoning-token-billing) maps them so the token comparison is fair.

If something breaks after the switch, the symptom usually points to one of the mistakes OpenAI lists in the migration guide:

| Symptom after switching | Likely cause |
|---|---|
| `AttributeError` or empty text where the answer should be | Code still reads `choices[0].message.content` |
| Crashes or missing text on reasoning or tool responses | Every `output` entry is treated as a message |
| Tool loop errors or incoherent follow-up turns | Reasoning, `function_call`, or `function_call_output` Items were dropped from the replayed context |
| Tool result rejected | `function_call_output` sent without the matching `call_id` |
| Structured output ignored or request rejected | `response_format` sent to Responses instead of `text.format` |
| Streaming UI shows nothing or never finishes | Chat chunk handler reused instead of branching on event types |
| Input token bill higher than expected | Assumed `previous_response_id` stops billing for earlier turns |

## FAQ

### Is OpenAI going to shut down Chat Completions?

There is no announced date. As of September 25, 2026, the deprecations page lists nothing for `/v1/chat/completions`, and the migration guide calls it supported. The API that shut down on August 26, 2026, was the Assistants API.

### Is the Responses API cheaper?

Not automatically. OpenAI reports 40% to 80% better cache utilization in its internal tests, which can lower cost if your prompts share cacheable prefixes. Chaining with `previous_response_id` doesn't reduce cost by itself, because prior input tokens are still billed.

### Can I pass my existing `messages` array to `client.responses.create`?

Yes, for simple role/content messages without function calls or multimodal inputs. Pass it as `input`. You still need to change how you read the result, and tool messages must be converted to `function_call` and `function_call_output` Items.

### Do messages have to alternate between user and assistant roles?

OpenAI's guides for Chat Completions and Responses don't document a rule that forces strict user/assistant alternation. What they do document is the mapping: system or developer guidance can move to `instructions`, and tool calls and results become separate Items linked by `call_id`. If you send the same transcript to other providers' chat APIs, check each provider's own message rules.

### Does Chat Completions work with GPT-6 Astra?

For plain text, yes; the model page lists Chat Completions as a supported endpoint. For function calling, no: OpenAI says Astra requires the Responses API for tool calling, and Astra rejects `reasoning_effort: "none"`.
