Inference Pricing

Last verified 16 Sep 2026

Inference provides a single control plane for managing inference workflows. It includes a Model Catalog where you can view available foundation models, including both DigitalOcean-hosted and third-party commercial models, compare model capabilities and pricing, use routing to match inference requests to the best-fit model, and run inference using serverless or dedicated deployments.

Inference has a usage-based pricing model, so costs scale with your actual usage.

Bring Your Own Models (BYOM)

BYOM model weights are stored in a service-managed, non-accessible Spaces location, and are billed at $5.00 per month. We do not charge you for browsing or managing imported models in Model Catalog. Costs apply only for storing model weights and for using those models with other paid features, such as dedicated inference deployments.

Model Playground

Usage is charged at the same rate as serverless inference.

Serverless Inference

Serverless inference is billed by DigitalOcean for both open-source and commercial models. Prices align with each provider’s published rates for transparency.

Warning

Serverless inference is prepaid only. You must maintain a positive prepaid account balance to send serverless inference requests, and we deduct usage charges from this balance. If your balance reaches $0, access is suspended until you replenish it. To add a balance or enable auto-reload, see Manage Serverless Inference Prepayment.

The following shows pricing for foundation models available through serverless inference.

Anthropic Models
Note

When using Anthropic commercial models with your own model API keys, billing is handled directly by Anthropic at the provider’s rates.

Claude Sonnet 5 and Sonnet 4.5 support an input context window of up to 1M tokens.

Model Serverless Inference
USD per 1M tokens unless noted
Claude Fable 5.1 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$10.00$50.00$12.50$20.00$0.25
Claude Fable 5 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$10.00$50.00$12.50$20.00$1.00
Claude Haiku 4.5 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$1.00$5.00$1.25$2.00$0.10
Claude Opus 5 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$5.00$25.00$6.25$10.00$0.50Fast ModeAll prompts$10.00$50.00$12.50$20.00$1.00
Claude Opus 4.8 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$5.00$25.00$6.25$10.00$0.50Fast ModeAll prompts$10.00$50.00$12.50$20.00$1.00
Claude Opus 4.7 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$5.00$25.00$6.25$10.00$0.50
Claude Opus 4.6 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$5.00$25.00$6.25$10.00$0.50
Claude Opus 4.5 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$5.00$25.00$6.25$10.00$0.50
Claude Sonnet 5 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$2.00$10.00$2.50$4.00$0.20
Claude Sonnet 4.6 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandardAll prompts$3.00$15.00$3.75$6.00$0.30
Claude Sonnet 4.5 Processing modePrompt lengthInputOutputCache write (5m)Cache write (1h)Cache readStandard≤ 200K tokens$3.00$15.00$3.75$6.00$0.30Standard> 200K tokens$6.00$22.50N/AN/AN/A
Arcee Models
Model Serverless Inference
USD per 1M tokens unless noted
Trinity Large Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.25$0.90$0.06
fal Models
Model Serverless Inference
USD per 1M tokens unless noted
Fast SDXL $0.0011 per compute second
Flux Schnell $0.0030 per megapixel
Stable Audio 2.5 (Text-to-Audio) $0.00058 per compute second
Multilingual TTS v2 $0.10 per 1000 characters
OpenAI Models
Note

When using OpenAI commercial models with your own model API keys, billing is handled directly by OpenAI at the provider’s rates.

GPT-6 Astra, GPT-5.6 Sol, Terra, and Luna support an input context window of up to 1.05M tokens. GPT-5.5 and GPT-5.4 support an input context window of up to 1M tokens. GPT-6 Astra, GPT-5.6 Sol, Terra, Luna, GPT-5.5, GPT-5.4, and GPT-5.4 pro use long-context pricing for prompts over 272K tokens.

Model Serverless Inference
USD per 1M tokens unless noted
gpt-oss-120b Processing modePrompt lengthInputOutputStandardAll prompts$0.055$0.385
gpt-oss-20b Processing modePrompt lengthInputOutputStandardAll prompts$0.05$0.45
GPT-6 Astra Processing modePrompt lengthInputOutputCache writeCache readStandard≤ 272K tokens$10.00$50.00$12.50$1.00Standard> 272K tokens$20.00$75.00$25.00$2.00Fast Mode≤ 272K tokens$20.00$100.00$25.00$2.00Fast Mode> 272K tokens$40.00$150.00$50.00$4.00Flex Mode≤ 272K tokens$5.00$25.00$6.25$0.50Flex Mode> 272K tokens$10.00$37.50$12.50$1.00
GPT-5.6 Sol Processing modePrompt lengthInputOutputCache writeCache readStandard≤ 272K tokens$4.00$20.00$5.00$0.40Standard> 272K tokens$8.00$30.00$10.00$0.80Fast Mode≤ 272K tokens$8.00$40.00$10.00$0.80Fast Mode> 272K tokens$16.00$60.00$20.00$1.60Flex Mode≤ 272K tokens$2.00$10.00$2.50$0.20Flex Mode> 272K tokens$4.00$15.00$5.00$0.40
GPT-5.6 Terra Processing modePrompt lengthInputOutputCache writeCache readStandard≤ 272K tokens$2.00$12.00$2.50$0.20Standard> 272K tokens$4.00$18.00$5.00$0.40Fast Mode≤ 272K tokens$4.00$24.00$5.00$0.40Fast Mode> 272K tokens$8.00$36.00$10.00$0.80Flex Mode≤ 272K tokens$1.00$6.00$1.25$0.10Flex Mode> 272K tokens$2.00$9.00$2.50$0.20
GPT-5.6 Luna Processing modePrompt lengthInputOutputCache writeCache readStandard≤ 272K tokens$0.20$1.20$0.25$0.02Standard> 272K tokens$0.40$1.80$0.50$0.04Fast Mode≤ 272K tokens$0.40$2.40$0.50$0.04Fast Mode> 272K tokens$0.80$3.60$1.00$0.08Flex Mode≤ 272K tokens$0.10$0.60$0.125$0.01Flex Mode> 272K tokens$0.20$0.90$0.25$0.02
GPT-5.5 Processing modePrompt lengthInputOutputCache readStandard≤ 272K tokens$5.00$30.00$0.50Standard> 272K tokens$10.00$45.00$1.00Fast Mode≤ 272K tokens$12.50$75.00$1.25Flex Mode≤ 272K tokens$2.50$15.00$0.25Flex Mode> 272K tokens$5.00$22.50$0.50
GPT-5.4 Processing modePrompt lengthInputOutputCache readStandard≤ 272K tokens$2.50$15.00$0.25Standard> 272K tokens$5.00$22.50$0.50Fast Mode≤ 272K tokens$5.00$30.00$0.50Flex Mode≤ 272K tokens$1.25$7.50$0.13Flex Mode> 272K tokens$2.50$11.25$0.25
GPT-5.4 mini Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.75$4.50$0.075Fast ModeAll prompts$1.50$9.00$0.15Flex ModeAll prompts$0.375$2.25$0.0375
GPT-5.4 nano Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.20$1.25$0.02Flex ModeAll prompts$0.10$0.625$0.01
GPT-5.4 pro Processing modePrompt lengthInputOutputStandard≤ 272K tokens$30.00$180.00Standard> 272K tokens$60.00$270.00Flex Mode≤ 272K tokens$15.00$90.00Flex Mode> 272K tokens$30.00$135.00
GPT-5.3-Codex Processing modePrompt lengthInputOutputCache readStandardAll prompts$1.75$14.00$0.175Fast ModeAll prompts$3.50$28.00$0.35
GPT-5.2 Processing modePrompt lengthInputOutputCache readStandardAll prompts$1.75$14.00$0.175Fast ModeAll prompts$3.50$28.00$0.35Flex ModeAll prompts$0.875$7.00$0.0875
GPT-5.2 pro Processing modePrompt lengthInputOutputStandardAll prompts$21.00$168.00
GPT-5 Processing modePrompt lengthInputOutputCache readStandardAll prompts$1.25$10.00$0.125Fast ModeAll prompts$2.50$20.00$0.25Flex ModeAll prompts$0.625$5.00$0.0625
GPT-5 mini Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.25$2.00$0.025Fast ModeAll prompts$0.45$3.60$0.045Flex ModeAll prompts$0.125$1.00$0.0125
GPT-5 nano Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.05$0.40$0.005Flex ModeAll prompts$0.025$0.20$0.0025
GPT-4.1 Processing modePrompt lengthInputOutputCache readStandardAll prompts$2.00$8.00$0.50Fast ModeAll prompts$3.50$14.00$0.875
GPT-4o Processing modePrompt lengthInputOutputCache readStandardAll prompts$2.50$10.00$1.25Fast ModeAll prompts$4.25$17.00$2.125
GPT-4o mini Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.15$0.60$0.075Fast ModeAll prompts$0.25$1.00$0.125
o1 Processing modePrompt lengthInputOutputCache readStandardAll prompts$15.00$60.00$7.50
o3 Processing modePrompt lengthInputOutputCache readStandardAll prompts$2.00$8.00$0.50Fast ModeAll prompts$3.50$14.00$0.875Flex ModeAll prompts$1.00$4.00$0.25
o3-mini Processing modePrompt lengthInputOutputCache readStandardAll prompts$1.10$4.40$0.55
GPT-image-1 Processing modePrompt lengthInputOutputCache readStandardAll prompts$5.00$40.00$1.25
GPT Image 1.5 Processing modePrompt lengthInputOutputCache readStandardAll prompts$5.00$10.00$1.00
GPT Image 2 Text input$5.00 per 1M tokens
Text output$0.00 per 1M tokens
Text cache read$1.25 per 1M tokens
Image input$8.00 per 1M tokens
Image output$30.00 per 1M tokens
Image cache read$2.00 per 1M tokens
DigitalOcean-Hosted Models
Provider Model Serverless Inference
USD per 1M tokens unless noted
Alibaba Qwen3.8-Max Processing modePrompt lengthInputOutputCache readStandardAll prompts$2.00$6.00$0.20
Alibaba Qwen 3.5 397B A17B Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.55$3.50$0.111
Alibaba Qwen 3 TTS (1.7B) $20.00 per 1M character tokens
Alibaba Wan2.2-T2V-A14B $0.60 per video
DeepSeek DeepSeek V4.1 Flash Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.30$1.20$0.006
DeepSeek DeepSeek V4 Pro 0813 Processing modePrompt lengthInputOutputCache readStandardAll prompts$1.32$3.96$0.044
DeepSeek DeepSeek V4 Flash 0731 Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.08$0.252$0.0252
DeepSeek DeepSeek V4 Pro Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.87$1.74$0.174
DeepSeek DeepSeek V4 Flash Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.0679$0.168$0.0168
DeepSeek DeepSeek V3.2 Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.25$0.80$0.075
Google Gemma 4 Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.18$0.50$0.036
MiniMax MiniMax M2.5 (Public Preview) Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.30$1.20$0.06
Moonshot AI Kimi K3 Processing modePrompt lengthInputOutputCache readStandardAll prompts$2.55$12.95$0.285
Moonshot AI Kimi K2.6 Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.95$4.00$0.19
Meta Llama 4 Maverick 17B 128E Instruct Processing modePrompt lengthInputOutputStandardAll prompts$0.20$0.696
Mistral AI Ministral 3 14B Instruct Processing modePrompt lengthInputOutputStandardAll prompts$0.20$0.20
NVIDIA Nemotron 3 Ultra Processing modePrompt lengthInputOutputStandardAll prompts$0.90$1.70
NVIDIA Nemotron Nano 3 Omni Processing modePrompt lengthInputOutputStandardAll prompts$0.50$0.90
NVIDIA Nemotron Nano 12B v2 VL Processing modePrompt lengthInputOutputStandardAll prompts$0.20$0.60
Stability AI Stable Diffusion 3.5 Large $0.08 per image
Xiaomi MiMo V2.5 Pro Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.40$1.50$0.08
Z.ai GLM-5.3 Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.95$3.40$0.20
Z.ai GLM-5.3 Flash Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.15$0.50$0.03
Z.ai GLM-5.2 Processing modePrompt lengthInputOutputCache readStandardAll prompts$0.70$2.20$0.105

Dedicated Inference

Dedicated Inference is billed per GPU-hour based on the GPU you use.

GPU Price
AMD MI300X $2.59 per hour
AMD MI300X (8x) $20.70 per hour
AMD MI325X $2.98 per hour
AMD MI325X (8x) $23.82 per hour
AMD MI350X $6.89 per hour
NVIDIA B300 $10.39 per hour
NVIDIA B300 (8x) $83.10 per hour
NVIDIA H100 $4.41 per hour
NVIDIA H100 (8x) $30.32 per hour
NVIDIA H200 $4.47 per hour
NVIDIA H200 (8x) $35.78 per hour

Batch Inference

Batch inference is charged at up to a 50% discount on OpenAI and Anthropic models.

You are only charged for completed requests. If a batch job fails, is blocked by guardrails, or expires partway through, requests that were not processed are not charged.

Inference Router public

Inference Router is available in public preview and enabled for all users. You can contact support for questions or assistance.

There is no additional cost to using Inference Router during public preview. Using inference routing forwards requests to foundation models for serverless inference. You are billed for the models that serve each request.

Tools Usage

Knowledge base retrieval, DigitalOcean MCP servers, and Anthropic- and OpenAI-only tools, such as tool search and computer use, do not incur additional charges other than the standard per-token inference costs.

The following tools incur charges in addition to the standard per-token inference costs:

  • Web search: $10 per 1000 requests, not charged when using Anthropic models
  • Web fetch: $3 per 1000 requests, not charged when using Anthropic models

The model synthesis tool does not add a separate charge. One model synthesis tool request can invoke multiple panel models, the primary judge model, and web search or web fetch calls. Usage and billing reflect all of that work at the existing per-model token rates and web tool rates above.

DigitalOcean Evaluations

DigitalOcean Evaluations can use a model or an Inference Router configuration as the candidate. Evaluations that use candidate models deployed on Serverless Inference, or judge models, are charged at the same token rates as serverless inference.

Candidate models deployed on Dedicated Inference do not incur additional evaluation-specific token charges.

Storage for evaluation datasets and evaluation results is currently provided at no additional charge. DigitalOcean may introduce or modify storage fees in the future.

Knowledge Bases

Knowledge base pricing is shown per million tokens, but billing is calculated per thousand tokens.

You’re billed for both indexing and storage:

  • Tokens used for indexing and retrieval query vectorization: We charge for tokens used to generate embeddings during indexing and to vectorize user queries during retrieval. Both use the same embeddings model pricing.

    Indexing pricing is the same for manual and auto-indexing. Indexing charges apply only when changes are detected, such as new, updated, or deleted files or URLs. If auto-indexing is paused or no changes are found, there are no indexing charges.

    Note

    Retrieval requests sent through a MCP server are billed the same as retrieval requests sent directly to the knowledge base retrieve endpoint. This includes the tokens used to vectorize the retrieval query with the selected embeddings model.

    For example, a 10 MB dataset is about 3 million tokens, and a 1 GB dataset is about 250 million tokens.

    Actual costs depend on the embeddings model:

    Model Price
    all-mini-lm-l6-v2 $0.009 per 1M input tokens
    multi-qa-mpnet-base-dot-v1 $0.009 per 1M input tokens
    gte-large-en-v1.5 $0.09 per 1M input tokens
    Qwen3 Embedding 0.6B $0.04 per 1,000,000 tokens
    BGE-M3 $0.02 per 1,000,000 tokens
    E5 Large V2 $0.02 per 1,000,000 tokens
    Note

    One token is roughly four characters (approximately 75 words per 100 tokens). Non-Latin scripts, emojis, or binary data may increase token counts.

  • Reranking tokens: If reranking is enabled, tokens used to rerank results are billed based on the selected reranking model. For supported reranking models, see available reranking models.

    Model Price
    BGE Reranker v2 m3 $0.01 per 1M reranking tokens
  • Storage: Embeddings are stored in OpenSearch. See OpenSearch pricing.

Chunking has no separate charge. Chunking costs depend on embedding token usage, OpenSearch database, and the selected embeddings model.

Chunking strategy cost depends on how many tokens the strategy embeds and returns:

  • Section-based and fixed length chunking are the most cost-efficient because they use simple splitting and have predictable token usage.
  • Semantic chunking costs more because it uses the embeddings model to detect semantic boundaries and embed final chunks, often resulting in 1.5 to 3 times more indexing tokens.
  • Hierarchical chunking slightly increases indexing cost by creating parent and child embeddings. It can also increase retrieval cost because agents receive both child and parent chunks for each lookup.

Changing your chunking strategy or configuration requires re-indexing the affected data source, which consumes additional tokens. For guidance on chunking configurations and best practices, see our chunking parameters reference and chunking best practices.

If you use RAG Playground, answer generation is billed separately based on the selected serverless inference model. Free tokens for RAG Playground are not separate; they are shared with Model Playground.

Agent Platform

Agent creation is free. We charge for model usage and for additional features like knowledge bases and guardrails. We display prices per million tokens and bill per thousand tokens for accuracy.

Model usage is billed by DigitalOcean. You are charged for all input and output tokens processed by the agent at the same token rates as serverless inference. Token usage depends on factors such as input length, agent instructions, attached knowledge bases, and configuration settings. To optimize usage, test your agents and adjust their parameters.

Agent Guardrails

Charges apply for all tokens processed through agent guardrails:

Guardrail Price
Content Moderation $0.20 per 1,000,000 tokens
Jailbreak Detection $0.20 per 1,000,000 tokens
Sensitive Data Detection $0.34 per 1,000,000 tokens

Costs are per token. Creating, editing, or duplicating guardrails has no additional cost.

Functions

If you attach DigitalOcean Functions to your agent, you are billed at functions pricing.

Agent Evaluations

Agent evaluations are charged by token usage at the same rates as model usage.

Agent Development Kit public

You are not charged for using the Agent Development Kit during public preview. However, you are billed for other DigitalOcean Inference features you use with your agent deployment:

  • We charge for model usage for Agent Development Kit (ADK). If you are using a DigitalOcean-hosted model, you are charged for those model keys.
Note

For General Availability, agent deployment hosting, measured in GiB-sec, will be charged. We will also be charging for judge input and output tokens, which are the tokens used for judging the agent inputs and outputs against the test case’s chosen metrics. These costs are waived during public preview.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.