推理 API

Chat Completions

查看 Markdown

Chat Completions API 是 Responses API 的前身,采用无状态设计并兼容 OpenAI。新集成应使用 Responses;请参阅从 Chat Completions 迁移



聊天补全

/v1/chat/completions

根据文本或图像聊天提示创建聊天响应。这是向聊天和图像理解模型发送请求的端点。

请求体

响应体

choicesarray<object>

模型的响应候选列表。长度对应请求体中的 `n`(默认为 1)。

createdinteger

聊天补全创建时间,以 Unix 时间戳表示。

idstring

聊天响应的唯一 ID。

modelstring

创建聊天补全所用的模型 ID。

objectstring

对象类型,始终为 `"chat.completion"`。

service_tier"default" | "priority" | "fast"

("default" | "priority" | "fast",必填)— 请求的处理层级。`Fast` 和 `Priority` 可互换:若模型配置了快速部署,则使用该部署及其费率,否则以更高价格提供更高调度优先级。

Exampletext

text

{
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful assistant that can answer questions and help with tasks."
    },
    {
      "role": "user",
      "content": "What is 101*3?"
    }
  ],
  "model": "latest"
}
Exampletext

text

{
  "id": "a3d1008e-4544-40d4-d075-11527e794e4a",
  "object": "chat.completion",
  "created": 1752854522,
  "model": "latest",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "101 multiplied by 3 is 303.",
        "refusal": null
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 32,
    "completion_tokens": 9,
    "total_tokens": 135,
    "prompt_tokens_details": {
      "text_tokens": 32,
      "audio_tokens": 0,
      "image_tokens": 0,
      "cached_tokens": 6
    },
    "completion_tokens_details": {
      "reasoning_tokens": 94,
      "audio_tokens": 0,
      "accepted_prediction_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "num_sources_used": 0
  },
  "system_fingerprint": "fp_3a7881249c"
}

获取延迟 chat completions

/v1/chat/deferred-completion/{request_id}

尝试获取先前启动的延迟补全的结果。请求已完成时,返回 `200 Success` 和响应体;请求等待处理时,返回 `202 Accepted`。

路径参数

request_idstring

先前延迟聊天请求返回的延迟请求 ID。

响应体

choicesarray<object>

模型的响应候选列表。长度对应请求体中的 `n`(默认为 1)。

createdinteger

聊天补全创建时间,以 Unix 时间戳表示。

idstring

聊天响应的唯一 ID。

modelstring

创建聊天补全所用的模型 ID。

objectstring

对象类型,始终为 `"chat.completion"`。

service_tier"default" | "priority" | "fast"

("default" | "priority" | "fast",必填)— 请求的处理层级。`Fast` 和 `Priority` 可互换:若模型配置了快速部署,则使用该部署及其费率,否则以更高价格提供更高调度优先级。

Exampletext

text

No parameters.
Exampletext

text

{
  "id": "335b92e4-afa5-48e7-b99c-b9a4eabc1c8e",
  "object": "chat.completion",
  "created": 1743770624,
  "model": "latest",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "101 multiplied by 3 is 303.",
        "refusal": null
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 31,
    "completion_tokens": 11,
    "total_tokens": 42,
    "prompt_tokens_details": {
      "text_tokens": 31,
      "audio_tokens": 0,
      "image_tokens": 0,
      "cached_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0,
      "audio_tokens": 0,
      "accepted_prediction_tokens": 0,
      "rejected_prediction_tokens": 0
    }
  },
  "system_fingerprint": "fp_156d35dcaa"
}

最后更新:2026 年 9 月 2 日