Release Note
Last verified 20 Aug 2026
Inference Router now uses cache-aware routing to maximize prompt cache reuse and automatically applies prompt caching to eligible Anthropic requests. You can opt out of prompt caching for supported models with X-Model-Affinity: none and control cache-aware model switching with the x-routing-max-switch-spend-pct header. For more information, see Use Prompt Caching.