x-ratelimit-reset-<limit-type> is not a fixed-window boundary. Since inference buckets refill continuously, the value is a refill projection.
Quota-Specific Response Headers For Serverless Inference
Last verified 21 Sep 2026
Inference provides a single control plane for managing inference workflows. It includes a Model Catalog where you can view available foundation models, including both DigitalOcean-hosted and third-party commercial models, compare model capabilities and pricing, use routing to match inference requests to the best-fit model, and run inference using serverless or dedicated deployments.
Every serverless inference response includes rate-limit headers. Token headers are scoped to the model you called, and request headers are scoped to the model’s request group:
x-ratelimit-limit-<limit-type>: Total capacity configured for the evaluated quota bucket, corresponding to the number of requests (x-ratelimit-limit-requests), tokens per day (x-ratelimit-limit-tokens-per-day), and tokens per minute (x-ratelimit-limit-tokens-per-minute).x-ratelimit-remaining-<limit-type>: Remaining available capacity after the current request evaluation, corresponding to the number of remaining requests (x-ratelimit-remaining-requests), remaining tokens per day (x-ratelimit-remaining-tokens-per-day), and remaining tokens per minute (x-ratelimit-remaining-tokens-per-minute).x-ratelimit-reset-<limit-type>: Unix epoch timestamp (a non-zero value in seconds) that projects when the bucket will be refilled to enough capacity to satisfy a request of the size that was just evaluated or rejected. Because the bucket refills continuously, this is a forward projection, not a window-boundary timestamp. These headers correspond to the number of requests (x-ratelimit-reset-requests), tokens per day (x-ratelimit-reset-tokens-per-day), and tokens per minute (x-ratelimit-reset-tokens-per-minute).
A value of 0 indicates that when the request was evaluated, there was sufficient capacity, and no action is required.
For each limit, the first header is your limit, the second is how much you have left right now, and the third is a Unix epoch timestamp for when that limit is projected to have enough capacity again. Wait max(reset_timestamp - now, 1s) seconds before retrying, where now is the current Unix time in seconds. The following example shows calling one of the larger models:
x-ratelimit-limit-requests: 1201
x-ratelimit-limit-tokens-per-day: 11000
x-ratelimit-limit-tokens-per-minute: 15000000
x-ratelimit-remaining-requests: 1200
x-ratelimit-remaining-tokens-per-day: 1100
x-ratelimit-remaining-tokens-per-minute: 15000000
x-ratelimit-reset-requests: 1789592460
x-ratelimit-reset-tokens-per-day: 0
x-ratelimit-reset-tokens-per-minute: 0Your code can read these headers and slow down before ever hitting a limit. 429 responses also include a Retry-After header with the number of seconds to wait before retrying.
For troubleshooting rate limiting, see the Inference support articles.