Release Note

Last verified 20 Aug 2026

Inference Router now uses cache-aware routing to maximize prompt cache reuse and automatically applies prompt caching to eligible Anthropic requests. You can opt out of prompt caching for supported models with X-Model-Affinity: none and control cache-aware model switching with the x-routing-max-switch-spend-pct header. For more information, see Use Prompt Caching.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.