Release Note

Last verified 1 Jul 2026

Prompt caching for open-source models in serverless inference chat completions and responses API is now in public preview. Open-source models cache context automatically, so you do not need to set the cache_control or prompt_cache_retention parameters.

Prompt caching is available for the following open-source models:

  • DeepSeek V3.2
  • DeepSeek V4 Pro
  • DeepSeek V4 Flash
  • Kimi K2.5
  • Kimi K2.6
  • GLM 5
  • GLM-5.1
  • GLM-5.2
  • gpt-oss-120b
  • MiMo V2.5
  • MiMo V2.5 Pro
  • MiniMax M2.5
  • Qwen 3.5
  • Qwen3 Coder Flash

For more information, see Use Prompt Caching.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.