Prompt Caching
哪些情况会破坏 Cache
对先前 message 的任何修改都会破坏 cache。只能在末尾追加新的 message。
对于 reasoning model,可以通过以下任一方式保持 cache hit:
回传 encrypted reasoning content — 包含上一次 response 中的
reasoning_content。详情请参阅Encrypted Reasoning Content。使用 Stateful Response — 使用
previous_response_id自动继续对话。详情请参阅串联对话。
Cache hit:追加新的 message
Prompt prefix 与上一次请求完全相同,只追加了一条新的 user message:
# Turn 1: Initial request (establishes the cache)
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.5",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."}
]
}'
# Turn 2: Cache HIT — exact prefix preserved, new message appended
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.5",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "user", "content": "Show me a code example."}
]
}'Cache miss:编辑先前的 message
修改任意一条先前 message 的内容都会破坏 prefix 匹配:
# Cache MISS — editing the assistant message content
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.5",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "assistant", "content": "It stores KV pairs."},
{"role": "user", "content": "Show me a code example."}
]
}'发生的变化: 第 11 行的 Assistant Response 被缩短为 "It stores KV pairs."(第 12 行)。
Cache miss:删除 message
从对话中删除任意 message 都会破坏 prefix:
# Cache MISS — the assistant message was removed
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.5",
"messages": [
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "user", "content": "What is prompt caching?"},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "user", "content": "Show me a code example."}
]
}'发生的变化: 第 11 行的 Assistant Message 被完全删除。
Cache miss:重新排列 message
改变 message 顺序也会破坏 prefix:
# Cache MISS — user and system messages are swapped
curl https://api.x.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "x-grok-conv-id: conv_abc123" \
-d '{
"model": "grok-4.5",
"messages": [
{"role": "user", "content": "What is prompt caching?"},
{"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
{"role": "assistant", "content": "Prompt caching stores KV pairs from unchanged prompt prefixes so they can be reused on subsequent requests. This makes responses faster and cheaper."},
{"role": "user", "content": "Show me a code example."}
]
}'发生的变化: 第 9 行和第 10 行被交换,user message 现在位于 system message 之前。