Use Responses for new work; keep Chat Completions where it works
Use the Responses API (/v1/responses) for anything new, and for any flow where a reasoning model has to call tools. Keep Chat Completions (/v1/chat/completions) where it already works for plain text or Structured Outputs, where you need several candidates per request, or where the same code must also run against providers that only speak Chat Completions.
Chat Completions is not deprecated. OpenAI's migration guide says it "remains supported" while Responses "is recommended for all new projects," and that you can move one user flow at a time. What changed in 2026 is the model side: starting with GPT-5.4, Chat Completions only allows tool calling when reasoning_effort is none, and GPT-6 Astra can't call functions through Chat Completions at all.
| Your situation | Use | Deciding reason |
|---|---|---|
| Function calling with reasoning on GPT-5.4 or later | Responses | Chat Completions allows tool calls only with reasoning_effort: "none" |
| Function calling on GPT-6 Astra | Responses | Astra requires Responses for tool calling |
| OpenAI-hosted tools (web search, file search, code interpreter, remote MCP, image generation, computer use) | Responses | Chat Completions has no native hosted tools |
| Server-side conversation state | Responses | previous_response_id or the Conversations API |
| One code path for OpenAI and "OpenAI-compatible" providers | Chat Completions, or both | Many providers expose Chat Completions; Responses support varies |
Several candidate answers per request (n > 1) | Chat Completions | Responses removed n |
| Any other new project on OpenAI models | Responses | OpenAI's recommendation for all new projects |
| A text-only or Structured Outputs flow that already works | Stay, migrate later | Supported, with no shutdown date |
Read the rows from the top and stop at the first one that fits your flow. The decision is per flow, not per app: if one flow needs reasoning plus tools, that flow moves even if the rest stays put.

Is Chat Completions deprecated?
No. As of September 25, 2026, OpenAI's deprecations page has no entry that retires /v1/chat/completions. The confusion comes from three different APIs with similar names:
| API | Endpoint | Status as of September 25, 2026 |
|---|---|---|
| Completions (legacy) | /v1/completions | The older prompt-in, text-out API. Model pages label it "Completions (legacy)." |
| Assistants API | /v1/assistants | Shut down on August 26, 2026. OpenAI names the Responses API and the Conversations API as its replacement. |
| Chat Completions | /v1/chat/completions | Supported. No deprecation entry. |
| Responses | /v1/responses | Recommended for all new projects. |
OpenAI gave Assistants users notice on August 26, 2025, and removed the API exactly one year later. If your code still calls /v1/assistants, it needs the Responses and Conversations APIs now, not later.
One more source of doubt: OpenAI's own Codex CLI dropped its chat/completions wire support in February 2026, and the announcement describes that protocol as one that "originated in the GPT-3.5 era." That decision covers Codex as a client. Custom providers in Codex had to switch from wire_api = "chat" to wire_api = "responses", but the API endpoint itself kept running. If that is what brought you here, the practical fix is in Connect Codex to a Third-Party Model API Without Guessing the Protocol.
The 2026 rule that decides most cases: tools plus reasoning
If your app calls tools, a model constraint can settle the question before any difference in API design does. OpenAI states it in three guides:
- The migration guide: "Starting with GPT-5.4, Chat Completions does not support tool calling with reasoning_effort values other than none."
- The reasoning guide and the function calling guide: "GPT-6 Astra requires the Responses API for tool calling." The reasoning guide adds that Astra rejects
nonereasoning effort with HTTP 400.
Put together, this is how tool calling looks per API for current models:
| Model | Chat Completions + tools | Responses + tools |
|---|---|---|
GPT-6 Astra (gpt-6-astra) | Not supported. Astra accepts low through max effort, never none | Supported, including hosted tools listed on the model page |
GPT-5.4 and later, for example GPT-5.6 Sol (gpt-5.6, alias of gpt-5.6-sol) or GPT-6 Sol (gpt-6-sol) | Only with reasoning_effort: "none" | Supported at any effort the model allows |
| Models before GPT-5.4 | Not restricted by this rule | Supported |
The middle row applies OpenAI's "starting with GPT-5.4" wording to the newer models; it is a reading of the rule, not a separate test. The GPT-5.6 Sol page lists medium as its default effort, so a Chat Completions request that sends tools to gpt-5.6 has to turn reasoning off. On Responses, the same request keeps reasoning and tools together. If you aren't sure which model family your app uses, OpenAI Text Models in 2026: GPT-5.6 Sol, Terra, and Luna Compared lists the current IDs.
OpenAI also claims quality and cost gains from Responses with reasoning models. From its internal evals: a 3% improvement on SWE-bench "with same prompt and setup," and 40% to 80% better cache utilization compared with Chat Completions. Those are vendor numbers, not an independent benchmark, and cache utilization doesn't translate into a guaranteed saving on your bill.
What changes in your code when you move to Responses
Most of the migration is a handful of renamed fields plus one real design change: output and input become arrays of typed Items (message, reasoning, function_call, function_call_output, and so on) instead of chat messages that bundle everything together.
| Concern | Chat Completions | Responses |
|---|---|---|
| Endpoint | POST /v1/chat/completions | POST /v1/responses |
| SDK call | client.chat.completions.create | client.responses.create |
| Input | messages | input (a string or an array of Items) |
| System or developer guidance | A system or developer message | Top-level instructions, or a compatible message Item |
| Where the text lives | choices[0].message.content | output_text (SDK helper) or the Items in output |
| Multiple candidates | n | Not available; make separate requests |
| Function definition | Externally tagged: {"type": "function", "function": {...}} | Internally tagged: {"type": "function", "name": ..., ...} |
| Function strictness | Non-strict unless you set strict | Tries strict by default, falls back to non-strict |
| Tool call and result | tool_calls on the assistant message, then a role: "tool" message with tool_call_id | function_call Item, then a function_call_output Item with the same call_id |
| Structured Outputs | response_format | text.format |
| Streaming | Chunks with a delta field | Typed server-sent events |
| Multi-turn state | You resend the whole messages array | previous_response_id, manual Item replay, or the Conversations API |
Basic text generation
A simple role/content array works as Responses input without changes, as long as it doesn't contain function calls or multimodal parts. The part that breaks is reading the result.
from openai import OpenAI
client = OpenAI()
messages = [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is a webhook?"},
]
# Chat Completions
completion = client.chat.completions.create(model="gpt-5.6", messages=messages)
print(completion.choices[0].message.content)
# Responses: same array as input
response = client.responses.create(model="gpt-5.6", input=messages)
print(response.output_text)
# Responses: guidance moved to instructions
response = client.responses.create(
model="gpt-5.6",
instructions="Answer in one sentence.",
input="What is a webhook?",
)
print(response.output_text)output_text joins the final text for you. When a response can contain reasoning or tool calls, loop over response.output and branch on each Item's type; the first Item is not always a message.
Function calling

The tool definition loses its nested function object. Also check strictness: Chat Completions functions are non-strict by default, while Responses attempts strict mode when you omit strict. If the schema can't be made strict, Responses falls back to best-effort calling and reports strict: false on the resolved tool. Set "strict": False yourself if you want the old behavior without surprises.
import json
# Chat Completions definition (externally tagged)
chat_tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
},
}]
# Responses definition (internally tagged)
tools = [{
"type": "function",
"name": "get_weather",
"description": "Get current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
}]
def get_weather(city: str) -> dict:
return {"city": city, "forecast": "light rain"} # replace with a real lookup
input_items = [{"role": "user", "content": "Do I need an umbrella in Paris today?"}]
response = client.responses.create(model="gpt-5.6", input=input_items, tools=tools)
# Keep every returned Item, including reasoning, before adding results
input_items += response.output
for item in response.output:
if item.type == "function_call":
args = json.loads(item.arguments)
input_items.append({
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(get_weather(**args)),
})
final = client.responses.create(model="gpt-5.6", input=input_items, tools=tools)
print(final.output_text)Two details matter with reasoning models. First, the function calling guide says reasoning Items returned alongside tool calls "must also be passed back with tool call outputs," which is why the loop appends all of response.output and not just the calls. Second, the model can return several calls in one turn. Handle each one, or set parallel_tool_calls to false to get zero or one call.
Structured Outputs
Both APIs support strict JSON Schema output. Only the location of the schema moves, from response_format to text.format, and the json_schema wrapper is flattened.
schema = {
"type": "object",
"properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
"required": ["name", "age"],
"additionalProperties": False,
}
# Chat Completions
client.chat.completions.create(
model="gpt-5.6",
messages=[{"role": "user", "content": "Jane, 54, joined in March."}],
response_format={
"type": "json_schema",
"json_schema": {"name": "person", "strict": True, "schema": schema},
},
)
# Responses
client.responses.create(
model="gpt-5.6",
input="Jane, 54, joined in March.",
text={"format": {"type": "json_schema", "name": "person", "strict": True, "schema": schema}},
)Sending response_format to /v1/responses is one of the mistakes OpenAI lists in its migration guide.
Streaming
A Chat Completions stream is a series of chunks with choices[0].delta. A Responses stream is a series of typed events, so the handler has to branch on event.type.
# Chat Completions
stream = client.chat.completions.create(model="gpt-5.6", messages=messages, stream=True)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
# Responses
stream = client.responses.create(model="gpt-5.6", input=messages, stream=True)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="")
elif event.type == "response.completed":
usage = event.response.usage
elif event.type == "error":
raise RuntimeError(event.message)The events OpenAI names for text streaming are response.created, response.output_text.delta, response.completed, and error. Function calls add response.function_call_arguments.delta and response.function_call_arguments.done. The streaming guide covers the other event types, such as response.failed and response.output_item.added.
Conversation state, storage, and billing
With Chat Completions, your application stores the transcript and sends the accumulated messages array every turn. The Responses API gives you three options:
previous_response_id: point the next request at the last response and OpenAI supplies the prior context.- Manual Item replay: pass earlier output Items back in
input, which lets you trim or edit context yourself. - The Conversations API: a persistent conversation object that lives until you delete it.
first = client.responses.create(
model="gpt-5.6",
instructions="You are a concise support agent.",
input="My invoice shows two charges for September.",
)
follow_up = client.responses.create(
model="gpt-5.6",
previous_response_id=first.id,
instructions="You are a concise support agent.", # not carried over, so resend it
input="Can you refund the duplicate?",
)Keep two limits in mind with previous_response_id. It doesn't carry over the previous response's top-level instructions, so resend them on every request. It also doesn't make earlier turns free: OpenAI states that "all previous input tokens for responses in the chain are billed as input tokens." Any cost benefit comes from caching, a separate mechanism covered in OpenAI vs Claude Prompt Caching Cost: Cached Token Pricing and Break-Even Math.
Storage defaults differ as well. According to OpenAI's data controls page:
- Responses are stored by default, with a 30-day application state retention period by default or when
storeistrue. Setstore: falseto turn that off. - Conversation objects are not subject to the 30-day limit; they stay until deleted, and
/v1/conversationsis not eligible for Zero Data Retention (ZDR). - For ZDR organizations,
storeis always treated asfalse. To keep reasoning context across turns without stored state, replay the encrypted reasoning Items the API returns (each carriesencrypted_content). - For
/v1/chat/completions, the data controls table lists no application state retention except audio outputs (1 hour). The migration guide, however, says "Chat completions are stored by default for new accounts."
Those two Chat Completions statements describe different kinds of storage. If you don't want outputs kept, set store: false explicitly on both APIs instead of relying on a default.
When staying on Chat Completions is the right call
Moving has a cost, and in several situations it buys you nothing yet:
- The flow is plain text or Structured Outputs, with no tools. It works, it's supported, and there's no shutdown date. Migrate it when you next touch that code.
- You need
n> 1. Responses returns one generation per request. Getting three candidates means three separate requests and your own fan-out logic. - Your code has to stay portable. Many third-party "OpenAI-compatible" endpoints implement Chat Completions; Responses support ranges from none to partial. Some providers accept Responses-style requests but ignore unsupported parameters and tools and keep no state. Check a named provider's docs before you assume parity.
- You use tools on GPT-5.4 or later and don't need reasoning for them. A simple extraction or routing call can run with
reasoning_effort: "none"on Chat Completions. The moment that flow needs reasoning, it has to move.
Staying doesn't mean freezing. OpenAI recommends "migrating all flows to the Responses API over time to take advantage of the latest OpenAI features and improvements," so keep the migration on your roadmap even for flows that can wait.
How to migrate one flow and prove it works
OpenAI's rollout checklist works best when you treat each step as a test that has to pass before more traffic moves. A workable order:
- Start with the simplest text flow. Change the endpoint and request body, then replace every read of
choices[0].message.contentwithoutput_textor an Item loop. - Pick one state strategy for that flow:
previous_response_id, manual Item replay, or Conversations. Resendinstructionsif you chain responses. - Set
storedeliberately. For stateless or ZDR flows, usestore: falseand replay encrypted reasoning Items. - Port function definitions and the tool loop. Verify that every
function_call_outputcarries the matchingcall_idand that reasoning Items travel with the results. Run a schema that used to be non-strict to see whether Responses accepts it strictly or falls back. - Move Structured Outputs from
response_formattotext.format. - Rewrite the stream handler for typed events, including
errorandresponse.failed. - Replace custom tool plumbing with hosted tools where they fit, such as web search or file search.
- Compare before shifting traffic. OpenAI's checklist ends with comparing behavior, latency, token usage, and errors between the two paths. Usage fields are named differently on each API; Reasoning Token Billing: Avoid Double Counting Across Gemini, OpenAI, and Claude maps them so the token comparison is fair.
If something breaks after the switch, the symptom usually points to one of the mistakes OpenAI lists in the migration guide:
| Symptom after switching | Likely cause |
|---|---|
AttributeError or empty text where the answer should be | Code still reads choices[0].message.content |
| Crashes or missing text on reasoning or tool responses | Every output entry is treated as a message |
| Tool loop errors or incoherent follow-up turns | Reasoning, function_call, or function_call_output Items were dropped from the replayed context |
| Tool result rejected | function_call_output sent without the matching call_id |
| Structured output ignored or request rejected | response_format sent to Responses instead of text.format |
| Streaming UI shows nothing or never finishes | Chat chunk handler reused instead of branching on event types |
| Input token bill higher than expected | Assumed previous_response_id stops billing for earlier turns |
FAQ
Is OpenAI going to shut down Chat Completions?
There is no announced date. As of September 25, 2026, the deprecations page lists nothing for /v1/chat/completions, and the migration guide calls it supported. The API that shut down on August 26, 2026, was the Assistants API.
Is the Responses API cheaper?
Not automatically. OpenAI reports 40% to 80% better cache utilization in its internal tests, which can lower cost if your prompts share cacheable prefixes. Chaining with previous_response_id doesn't reduce cost by itself, because prior input tokens are still billed.
Can I pass my existing messages array to client.responses.create?
Yes, for simple role/content messages without function calls or multimodal inputs. Pass it as input. You still need to change how you read the result, and tool messages must be converted to function_call and function_call_output Items.
Do messages have to alternate between user and assistant roles?
OpenAI's guides for Chat Completions and Responses don't document a rule that forces strict user/assistant alternation. What they do document is the mapping: system or developer guidance can move to instructions, and tool calls and results become separate Items linked by call_id. If you send the same transcript to other providers' chat APIs, check each provider's own message rules.
Does Chat Completions work with GPT-6 Astra?
For plain text, yes; the model page lists Chat Completions as a supported endpoint. For function calling, no: OpenAI says Astra requires the Responses API for tool calling, and Astra rejects reasoning_effort: "none".



