Inference API
Chat Completions
The Chat Completions API is the stateless, OpenAI-compatible predecessor of the Responses API. New integrations should use Responses; see Migrating from Chat Completions.
Chat completions
/v1/chat/completions
Create a chat response from text/image chat prompts. This is the endpoint for making requests to chat and image understanding models.
Request Body
Response Body
choicesarray<object>A list of response choices from the model. The length corresponds to the `n` in request body (default to 1).
createdintegerThe chat completion creation time in Unix timestamp.
idstringA unique ID for the chat response.
modelstringModel ID used to create chat completion.
objectstringThe object type, which is always `"chat.completion"`.
service_tier"default" | "priority" | "fast"Processing tier for a request. `Fast` and `Priority` are interchangeable: the model's fast deployment and its rates where configured, else higher scheduling priority at a higher price.
{
"messages": [
{
"role": "system",
"content": "You are a helpful assistant that can answer questions and help with tasks."
},
{
"role": "user",
"content": "What is 101*3?"
}
],
"model": "latest"
}{
"id": "a3d1008e-4544-40d4-d075-11527e794e4a",
"object": "chat.completion",
"created": 1752854522,
"model": "latest",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "101 multiplied by 3 is 303.",
"refusal": null
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 32,
"completion_tokens": 9,
"total_tokens": 135,
"prompt_tokens_details": {
"text_tokens": 32,
"audio_tokens": 0,
"image_tokens": 0,
"cached_tokens": 6
},
"completion_tokens_details": {
"reasoning_tokens": 94,
"audio_tokens": 0,
"accepted_prediction_tokens": 0,
"rejected_prediction_tokens": 0
},
"num_sources_used": 0
},
"system_fingerprint": "fp_3a7881249c"
}Get deferred chat completions
/v1/chat/deferred-completion/{request_id}
Tries to fetch a result for a previously-started deferred completion. Returns `200 Success` with the response body, if the request has been completed. Returns `202 Accepted` when the request is pending processing.
Path Parameters
request_idstringThe deferred request id returned by a previous deferred chat request.
Response Body
choicesarray<object>A list of response choices from the model. The length corresponds to the `n` in request body (default to 1).
createdintegerThe chat completion creation time in Unix timestamp.
idstringA unique ID for the chat response.
modelstringModel ID used to create chat completion.
objectstringThe object type, which is always `"chat.completion"`.
service_tier"default" | "priority" | "fast"Processing tier for a request. `Fast` and `Priority` are interchangeable: the model's fast deployment and its rates where configured, else higher scheduling priority at a higher price.
No parameters.{
"id": "335b92e4-afa5-48e7-b99c-b9a4eabc1c8e",
"object": "chat.completion",
"created": 1743770624,
"model": "latest",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "101 multiplied by 3 is 303.",
"refusal": null
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 31,
"completion_tokens": 11,
"total_tokens": 42,
"prompt_tokens_details": {
"text_tokens": 31,
"audio_tokens": 0,
"image_tokens": 0,
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0,
"audio_tokens": 0,
"accepted_prediction_tokens": 0,
"rejected_prediction_tokens": 0
}
},
"system_fingerprint": "fp_156d35dcaa"
}Last updated:September 2, 2026