Inference Reference

Last verified 1 Oct 2026

Inference provides a single control plane for managing inference workflows. It includes a Model Catalog where you can view available foundation models, including both DigitalOcean-hosted and third-party commercial models, compare model capabilities and pricing, use routing to match inference requests to the best-fit model, and run inference using serverless or dedicated deployments.

The DigitalOcean API

The DigitalOcean API lets you manage resources programmatically with standard HTTP requests. All actions available in the control panel are also available through the API.

  • Serverless Inference API: Interact directly with foundation models for chat completions, or generating image, audio and text-to-speech.

  • Dedicated Inference API: Manage your dedicated inference deployments. Dedicated Inference is available in public preview. You can opt in from the Feature Preview page.

  • Agent Platform API: Create, delete, and manage knowledge bases and generative AI agents. You can also use the API to add agent and function routes to agents, add data sources to knowledge bases, and start indexing jobs.

  • Agent Inference API: Interact with agents using an agent-specific endpoint.

The DigitalOcean Command Line Client, doctl

doctl is the command-line interface for the DigitalOcean API. It supports most of the same actions available in the API and DigitalOcean Control Panel.

  • Dedicated Inference: Manage your dedicated inference endpoints, including their accelerators, sizes, GPU model configuration, and access tokens. Dedicated Inference is available in public preview. You can opt in from the Feature Preview page.

  • Serverless Inference: Call models for chat completions, embeddings, images, messages, and responses, manage asynchronous invocations, list available models and regions, and manage OpenAI API keys.

  • Knowledge Base: Manage agent knowledge bases, attach and detach them from agents, and manage their data sources and indexing jobs.

See the doctl documentation or run a command with --help for more information.

The Inference SDK

Use the official DigitalOcean Python client library PyDo for:

As of 15 July 2026, the Gradient AI SDK has been deprecated. Use the official DigitalOcean TypeScript library or Go library.

The DigitalOcean MCP Server

The DigitalOcean MCP server lets you use natural language prompts to manage your DigitalOcean AI resources. You can:

  • Create, update, list, and delete Dedicated Inference endpoints
  • Interact with knowledge bases to retrieve relevant chunks, apply filters, and access indexed content for use in agent and retrieval workflows
  • Manage evaluation datasets, run evaluations and monitor agent deployments
  • Submit and retrieve batch inference jobs
  • Retrieve information from the model catalog.

All operations use argument-based input.

DigitalOcean MCP Servers

Use the DigitalOcean MCP server to manage your AI resources.

More Resources

Agent Evaluation Metrics

A list of available agent evaluation metrics and their definitions.

Chunking Parameters

Reference for DigitalOcean Knowledge Bases chunking parameters, their recommendations, and their constraints across supported embeddings models.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.