Skip to content

Chat Completions vs Responses API: Which One Should You Use?

Start new work and any tools-plus-reasoning flow on Responses. Keep Chat Completions for working text flows, n > 1, or providers that only speak Chat Completions.

•••11 min read•AI Model Comparison
Blue and pink blocks on a white platform labeled Chat Completions: Supported and Responses: Recommended, beside the title Chat Completions vs Responses API

Use Responses for new work; keep Chat Completions where it works

Use the Responses API (/v1/responses) for anything new, and for any flow where a reasoning model has to call tools. Keep Chat Completions (/v1/chat/completions) where it already works for plain text or Structured Outputs, where you need several candidates per request, or where the same code must also run against providers that only speak Chat Completions.

Chat Completions is not deprecated. OpenAI's migration guide says it "remains supported" while Responses "is recommended for all new projects," and that you can move one user flow at a time. What changed in 2026 is the model side: starting with GPT-5.4, Chat Completions only allows tool calling when reasoning_effort is none, and GPT-6 Astra can't call functions through Chat Completions at all.

Your situationUseDeciding reason
Function calling with reasoning on GPT-5.4 or laterResponsesChat Completions allows tool calls only with reasoning_effort: "none"
Function calling on GPT-6 AstraResponsesAstra requires Responses for tool calling
OpenAI-hosted tools (web search, file search, code interpreter, remote MCP, image generation, computer use)ResponsesChat Completions has no native hosted tools
Server-side conversation stateResponsesprevious_response_id or the Conversations API
One code path for OpenAI and "OpenAI-compatible" providersChat Completions, or bothMany providers expose Chat Completions; Responses support varies
Several candidate answers per request (n > 1)Chat CompletionsResponses removed n
Any other new project on OpenAI modelsResponsesOpenAI's recommendation for all new projects
A text-only or Structured Outputs flow that already worksStay, migrate laterSupported, with no shutdown date

Read the rows from the top and stop at the first one that fits your flow. The decision is per flow, not per app: if one flow needs reasoning plus tools, that flow moves even if the rest stays put.

Decision flow with four questions in order: tools with reasoning on GPT-5.4 or later or GPT-6 Astra leads to Responses; hosted tools or server-side state leads to Responses; a Chat-only provider or n greater than 1 leads to Chat Completions; a new flow leads to Responses; no to all four means keep Chat Completions and migrate later

Is Chat Completions deprecated?

No. As of September 25, 2026, OpenAI's deprecations page has no entry that retires /v1/chat/completions. The confusion comes from three different APIs with similar names:

APIEndpointStatus as of September 25, 2026
Completions (legacy)/v1/completionsThe older prompt-in, text-out API. Model pages label it "Completions (legacy)."
Assistants API/v1/assistantsShut down on August 26, 2026. OpenAI names the Responses API and the Conversations API as its replacement.
Chat Completions/v1/chat/completionsSupported. No deprecation entry.
Responses/v1/responsesRecommended for all new projects.

OpenAI gave Assistants users notice on August 26, 2025, and removed the API exactly one year later. If your code still calls /v1/assistants, it needs the Responses and Conversations APIs now, not later.

One more source of doubt: OpenAI's own Codex CLI dropped its chat/completions wire support in February 2026, and the announcement describes that protocol as one that "originated in the GPT-3.5 era." That decision covers Codex as a client. Custom providers in Codex had to switch from wire_api = "chat" to wire_api = "responses", but the API endpoint itself kept running. If that is what brought you here, the practical fix is in Connect Codex to a Third-Party Model API Without Guessing the Protocol.

The 2026 rule that decides most cases: tools plus reasoning

If your app calls tools, a model constraint can settle the question before any difference in API design does. OpenAI states it in three guides:

  • The migration guide: "Starting with GPT-5.4, Chat Completions does not support tool calling with reasoning_effort values other than none."
  • The reasoning guide and the function calling guide: "GPT-6 Astra requires the Responses API for tool calling." The reasoning guide adds that Astra rejects none reasoning effort with HTTP 400.

Put together, this is how tool calling looks per API for current models:

ModelChat Completions + toolsResponses + tools
GPT-6 Astra (gpt-6-astra)Not supported. Astra accepts low through max effort, never noneSupported, including hosted tools listed on the model page
GPT-5.4 and later, for example GPT-5.6 Sol (gpt-5.6, alias of gpt-5.6-sol) or GPT-6 Sol (gpt-6-sol)Only with reasoning_effort: "none"Supported at any effort the model allows
Models before GPT-5.4Not restricted by this ruleSupported

The middle row applies OpenAI's "starting with GPT-5.4" wording to the newer models; it is a reading of the rule, not a separate test. The GPT-5.6 Sol page lists medium as its default effort, so a Chat Completions request that sends tools to gpt-5.6 has to turn reasoning off. On Responses, the same request keeps reasoning and tools together. If you aren't sure which model family your app uses, OpenAI Text Models in 2026: GPT-5.6 Sol, Terra, and Luna Compared lists the current IDs.

OpenAI also claims quality and cost gains from Responses with reasoning models. From its internal evals: a 3% improvement on SWE-bench "with same prompt and setup," and 40% to 80% better cache utilization compared with Chat Completions. Those are vendor numbers, not an independent benchmark, and cache utilization doesn't translate into a guaranteed saving on your bill.

What changes in your code when you move to Responses

Most of the migration is a handful of renamed fields plus one real design change: output and input become arrays of typed Items (message, reasoning, function_call, function_call_output, and so on) instead of chat messages that bundle everything together.

ConcernChat CompletionsResponses
EndpointPOST /v1/chat/completionsPOST /v1/responses
SDK callclient.chat.completions.createclient.responses.create
Inputmessagesinput (a string or an array of Items)
System or developer guidanceA system or developer messageTop-level instructions, or a compatible message Item
Where the text liveschoices[0].message.contentoutput_text (SDK helper) or the Items in output
Multiple candidatesnNot available; make separate requests
Function definitionExternally tagged: {"type": "function", "function": {...}}Internally tagged: {"type": "function", "name": ..., ...}
Function strictnessNon-strict unless you set strictTries strict by default, falls back to non-strict
Tool call and resulttool_calls on the assistant message, then a role: "tool" message with tool_call_idfunction_call Item, then a function_call_output Item with the same call_id
Structured Outputsresponse_formattext.format
StreamingChunks with a delta fieldTyped server-sent events
Multi-turn stateYou resend the whole messages arrayprevious_response_id, manual Item replay, or the Conversations API

Basic text generation

A simple role/content array works as Responses input without changes, as long as it doesn't contain function calls or multimodal parts. The part that breaks is reading the result.

python
from openai import OpenAI

client = OpenAI()

messages = [
    {"role": "system", "content": "Answer in one sentence."},
    {"role": "user", "content": "What is a webhook?"},
]

# Chat Completions
completion = client.chat.completions.create(model="gpt-5.6", messages=messages)
print(completion.choices[0].message.content)

# Responses: same array as input
response = client.responses.create(model="gpt-5.6", input=messages)
print(response.output_text)

# Responses: guidance moved to instructions
response = client.responses.create(
    model="gpt-5.6",
    instructions="Answer in one sentence.",
    input="What is a webhook?",
)
print(response.output_text)

output_text joins the final text for you. When a response can contain reasoning or tool calls, loop over response.output and branch on each Item's type; the first Item is not always a message.

Function calling

Tool-call round trip in two lanes: Chat Completions sends messages with tools defined under function, gets message.tool_calls each with an id, and replies with a role tool message carrying tool_call_id; Responses sends input with tools named at the top level, gets reasoning and function_call Items with call_id, and sends back response.output plus function_call_output

The tool definition loses its nested function object. Also check strictness: Chat Completions functions are non-strict by default, while Responses attempts strict mode when you omit strict. If the schema can't be made strict, Responses falls back to best-effort calling and reports strict: false on the resolved tool. Set "strict": False yourself if you want the old behavior without surprises.

python
import json

# Chat Completions definition (externally tagged)
chat_tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    },
}]

# Responses definition (internally tagged)
tools = [{
    "type": "function",
    "name": "get_weather",
    "description": "Get current weather for a city.",
    "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
        "additionalProperties": False,
    },
}]

def get_weather(city: str) -> dict:
    return {"city": city, "forecast": "light rain"}  # replace with a real lookup

input_items = [{"role": "user", "content": "Do I need an umbrella in Paris today?"}]
response = client.responses.create(model="gpt-5.6", input=input_items, tools=tools)

# Keep every returned Item, including reasoning, before adding results
input_items += response.output
for item in response.output:
    if item.type == "function_call":
        args = json.loads(item.arguments)
        input_items.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(get_weather(**args)),
        })

final = client.responses.create(model="gpt-5.6", input=input_items, tools=tools)
print(final.output_text)

Two details matter with reasoning models. First, the function calling guide says reasoning Items returned alongside tool calls "must also be passed back with tool call outputs," which is why the loop appends all of response.output and not just the calls. Second, the model can return several calls in one turn. Handle each one, or set parallel_tool_calls to false to get zero or one call.

Structured Outputs

Both APIs support strict JSON Schema output. Only the location of the schema moves, from response_format to text.format, and the json_schema wrapper is flattened.

python
schema = {
    "type": "object",
    "properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
    "required": ["name", "age"],
    "additionalProperties": False,
}

# Chat Completions
client.chat.completions.create(
    model="gpt-5.6",
    messages=[{"role": "user", "content": "Jane, 54, joined in March."}],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "person", "strict": True, "schema": schema},
    },
)

# Responses
client.responses.create(
    model="gpt-5.6",
    input="Jane, 54, joined in March.",
    text={"format": {"type": "json_schema", "name": "person", "strict": True, "schema": schema}},
)

Sending response_format to /v1/responses is one of the mistakes OpenAI lists in its migration guide.

Streaming

A Chat Completions stream is a series of chunks with choices[0].delta. A Responses stream is a series of typed events, so the handler has to branch on event.type.

python
# Chat Completions
stream = client.chat.completions.create(model="gpt-5.6", messages=messages, stream=True)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

# Responses
stream = client.responses.create(model="gpt-5.6", input=messages, stream=True)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="")
    elif event.type == "response.completed":
        usage = event.response.usage
    elif event.type == "error":
        raise RuntimeError(event.message)

The events OpenAI names for text streaming are response.created, response.output_text.delta, response.completed, and error. Function calls add response.function_call_arguments.delta and response.function_call_arguments.done. The streaming guide covers the other event types, such as response.failed and response.output_item.added.

Conversation state, storage, and billing

With Chat Completions, your application stores the transcript and sends the accumulated messages array every turn. The Responses API gives you three options:

  1. previous_response_id: point the next request at the last response and OpenAI supplies the prior context.
  2. Manual Item replay: pass earlier output Items back in input, which lets you trim or edit context yourself.
  3. The Conversations API: a persistent conversation object that lives until you delete it.
python
first = client.responses.create(
    model="gpt-5.6",
    instructions="You are a concise support agent.",
    input="My invoice shows two charges for September.",
)

follow_up = client.responses.create(
    model="gpt-5.6",
    previous_response_id=first.id,
    instructions="You are a concise support agent.",  # not carried over, so resend it
    input="Can you refund the duplicate?",
)

Keep two limits in mind with previous_response_id. It doesn't carry over the previous response's top-level instructions, so resend them on every request. It also doesn't make earlier turns free: OpenAI states that "all previous input tokens for responses in the chain are billed as input tokens." Any cost benefit comes from caching, a separate mechanism covered in OpenAI vs Claude Prompt Caching Cost: Cached Token Pricing and Break-Even Math.

Storage defaults differ as well. According to OpenAI's data controls page:

  • Responses are stored by default, with a 30-day application state retention period by default or when store is true. Set store: false to turn that off.
  • Conversation objects are not subject to the 30-day limit; they stay until deleted, and /v1/conversations is not eligible for Zero Data Retention (ZDR).
  • For ZDR organizations, store is always treated as false. To keep reasoning context across turns without stored state, replay the encrypted reasoning Items the API returns (each carries encrypted_content).
  • For /v1/chat/completions, the data controls table lists no application state retention except audio outputs (1 hour). The migration guide, however, says "Chat completions are stored by default for new accounts."

Those two Chat Completions statements describe different kinds of storage. If you don't want outputs kept, set store: false explicitly on both APIs instead of relying on a default.

When staying on Chat Completions is the right call

Moving has a cost, and in several situations it buys you nothing yet:

  • The flow is plain text or Structured Outputs, with no tools. It works, it's supported, and there's no shutdown date. Migrate it when you next touch that code.
  • You need n > 1. Responses returns one generation per request. Getting three candidates means three separate requests and your own fan-out logic.
  • Your code has to stay portable. Many third-party "OpenAI-compatible" endpoints implement Chat Completions; Responses support ranges from none to partial. Some providers accept Responses-style requests but ignore unsupported parameters and tools and keep no state. Check a named provider's docs before you assume parity.
  • You use tools on GPT-5.4 or later and don't need reasoning for them. A simple extraction or routing call can run with reasoning_effort: "none" on Chat Completions. The moment that flow needs reasoning, it has to move.

Staying doesn't mean freezing. OpenAI recommends "migrating all flows to the Responses API over time to take advantage of the latest OpenAI features and improvements," so keep the migration on your roadmap even for flows that can wait.

How to migrate one flow and prove it works

OpenAI's rollout checklist works best when you treat each step as a test that has to pass before more traffic moves. A workable order:

  1. Start with the simplest text flow. Change the endpoint and request body, then replace every read of choices[0].message.content with output_text or an Item loop.
  2. Pick one state strategy for that flow: previous_response_id, manual Item replay, or Conversations. Resend instructions if you chain responses.
  3. Set store deliberately. For stateless or ZDR flows, use store: false and replay encrypted reasoning Items.
  4. Port function definitions and the tool loop. Verify that every function_call_output carries the matching call_id and that reasoning Items travel with the results. Run a schema that used to be non-strict to see whether Responses accepts it strictly or falls back.
  5. Move Structured Outputs from response_format to text.format.
  6. Rewrite the stream handler for typed events, including error and response.failed.
  7. Replace custom tool plumbing with hosted tools where they fit, such as web search or file search.
  8. Compare before shifting traffic. OpenAI's checklist ends with comparing behavior, latency, token usage, and errors between the two paths. Usage fields are named differently on each API; Reasoning Token Billing: Avoid Double Counting Across Gemini, OpenAI, and Claude maps them so the token comparison is fair.

If something breaks after the switch, the symptom usually points to one of the mistakes OpenAI lists in the migration guide:

Symptom after switchingLikely cause
AttributeError or empty text where the answer should beCode still reads choices[0].message.content
Crashes or missing text on reasoning or tool responsesEvery output entry is treated as a message
Tool loop errors or incoherent follow-up turnsReasoning, function_call, or function_call_output Items were dropped from the replayed context
Tool result rejectedfunction_call_output sent without the matching call_id
Structured output ignored or request rejectedresponse_format sent to Responses instead of text.format
Streaming UI shows nothing or never finishesChat chunk handler reused instead of branching on event types
Input token bill higher than expectedAssumed previous_response_id stops billing for earlier turns

FAQ

Is OpenAI going to shut down Chat Completions?

There is no announced date. As of September 25, 2026, the deprecations page lists nothing for /v1/chat/completions, and the migration guide calls it supported. The API that shut down on August 26, 2026, was the Assistants API.

Is the Responses API cheaper?

Not automatically. OpenAI reports 40% to 80% better cache utilization in its internal tests, which can lower cost if your prompts share cacheable prefixes. Chaining with previous_response_id doesn't reduce cost by itself, because prior input tokens are still billed.

Can I pass my existing messages array to client.responses.create?

Yes, for simple role/content messages without function calls or multimodal inputs. Pass it as input. You still need to change how you read the result, and tool messages must be converted to function_call and function_call_output Items.

Do messages have to alternate between user and assistant roles?

OpenAI's guides for Chat Completions and Responses don't document a rule that forces strict user/assistant alternation. What they do document is the mapping: system or developer guidance can move to instructions, and tool calls and results become separate Items linked by call_id. If you send the same transcript to other providers' chat APIs, check each provider's own message rules.

Does Chat Completions work with GPT-6 Astra?

For plain text, yes; the model page lists Chat Completions as a supported endpoint. For function calling, no: OpenAI says Astra requires the Responses API for tool calling, and Astra rejects reasoning_effort: "none".

Found an error, or hit a problem this article does not cover? Send the article URL and what you saw to hi@laozhang.ai. We check it and update the article.