Release Note
Last verified 1 Jul 2026
Prompt caching for open-source models in serverless inference chat completions and responses API is now in public preview. Open-source models cache context automatically, so you do not need to set the cache_control or prompt_cache_retention parameters.
Prompt caching is available for the following open-source models:
- DeepSeek V3.2
- DeepSeek V4 Pro
- DeepSeek V4 Flash
- Kimi K2.5
- Kimi K2.6
- GLM 5
- GLM-5.1
- GLM-5.2
- gpt-oss-120b
- MiMo V2.5
- MiMo V2.5 Pro
- MiniMax M2.5
- Qwen 3.5
- Qwen3 Coder Flash
For more information, see Use Prompt Caching.