prompt 缓存
哪些情况会破坏缓存
对先前消息的任何修改都会破坏缓存。只能在末尾追加新消息。
对于推理模型,可通过以下任一方式保持缓存命中:
回传加密的推理内容 — 包含上一次响应中的
reasoning_content。详情请参阅加密的推理内容。使用有状态响应 — 使用
previous_response_id自动继续对话。详情请参阅串联对话。
缓存命中 — 追加新消息
prompt 前缀与上一次请求完全相同,只追加了一条新的用户消息:
# Turn 1: Initial request (establishes the cache)
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.7",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."}
]
}'
# Turn 2: Cache HIT — exact prefix preserved, new message appended
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.7",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "user", "content": "Show me a code example."}
]
}'缓存未命中 — 编辑先前消息
修改任意一条先前消息的内容都会破坏前缀匹配:
# Cache MISS — editing the assistant message content
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.7",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "assistant", "content": "It stores KV pairs."},
{"role": "user", "content": "Show me a code example."}
]
}'发生的变化: 第 11 行的助手响应被缩短为 "It stores KV pairs."(第 12 行)。
缓存未命中 — 删除消息
从对话中删除任意消息都会破坏前缀:
# Cache MISS — the assistant message was removed
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.7",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "user", "content": "Show me a code example."}
]
}'发生的变化: 第 11 行的助手消息被完全删除。
缓存未命中 — 重新排列消息
改变消息顺序也会破坏前缀:
# Cache MISS — user and system messages are swapped
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.7",
"messages": [
{"role": "user", "content": "What is prompt caching?"},
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "user", "content": "Show me a code example."}
]
}'发生的变化: 第 9 行和第 10 行被交换,用户消息现在位于 system 消息之前。
后续内容
最后更新:2026 年 3 月 16 日