Inference API

Chat Completions

View as Markdown

The Chat Completions API is the stateless, OpenAI-compatible predecessor of the Responses API. New integrations should use Responses; see Migrating from Chat Completions.



Chat completions

/v1/chat/completions

Create a chat response from text/image chat prompts. This is the endpoint for making requests to chat and image understanding models.

Request Body

Response Body

choicesarray<object>

A list of response choices from the model. The length corresponds to the `n` in request body (default to 1).

createdinteger

The chat completion creation time in Unix timestamp.

idstring

A unique ID for the chat response.

modelstring

Model ID used to create chat completion.

objectstring

The object type, which is always `"chat.completion"`.

service_tier"default" | "priority" | "fast"

Processing tier for a request. `Fast` and `Priority` are interchangeable: the model's fast deployment and its rates where configured, else higher scheduling priority at a higher price.

Exampletext

text

{
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful assistant that can answer questions and help with tasks."
    },
    {
      "role": "user",
      "content": "What is 101*3?"
    }
  ],
  "model": "latest"
}
Exampletext

text

{
  "id": "a3d1008e-4544-40d4-d075-11527e794e4a",
  "object": "chat.completion",
  "created": 1752854522,
  "model": "latest",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "101 multiplied by 3 is 303.",
        "refusal": null
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 32,
    "completion_tokens": 9,
    "total_tokens": 135,
    "prompt_tokens_details": {
      "text_tokens": 32,
      "audio_tokens": 0,
      "image_tokens": 0,
      "cached_tokens": 6
    },
    "completion_tokens_details": {
      "reasoning_tokens": 94,
      "audio_tokens": 0,
      "accepted_prediction_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "num_sources_used": 0
  },
  "system_fingerprint": "fp_3a7881249c"
}

Get deferred chat completions

/v1/chat/deferred-completion/{request_id}

Tries to fetch a result for a previously-started deferred completion. Returns `200 Success` with the response body, if the request has been completed. Returns `202 Accepted` when the request is pending processing.

Path Parameters

request_idstring

The deferred request id returned by a previous deferred chat request.

Response Body

choicesarray<object>

A list of response choices from the model. The length corresponds to the `n` in request body (default to 1).

createdinteger

The chat completion creation time in Unix timestamp.

idstring

A unique ID for the chat response.

modelstring

Model ID used to create chat completion.

objectstring

The object type, which is always `"chat.completion"`.

service_tier"default" | "priority" | "fast"

Processing tier for a request. `Fast` and `Priority` are interchangeable: the model's fast deployment and its rates where configured, else higher scheduling priority at a higher price.

Exampletext

text

No parameters.
Exampletext

text

{
  "id": "335b92e4-afa5-48e7-b99c-b9a4eabc1c8e",
  "object": "chat.completion",
  "created": 1743770624,
  "model": "latest",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "101 multiplied by 3 is 303.",
        "refusal": null
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 31,
    "completion_tokens": 11,
    "total_tokens": 42,
    "prompt_tokens_details": {
      "text_tokens": 31,
      "audio_tokens": 0,
      "image_tokens": 0,
      "cached_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0,
      "audio_tokens": 0,
      "accepted_prediction_tokens": 0,
      "rejected_prediction_tokens": 0
    }
  },
  "system_fingerprint": "fp_156d35dcaa"
}

Last updated:September 2, 2026