Inference CLI Reference

Last verified 30 Sep 2026

Inference provides a single control plane for managing inference workflows. It includes a Model Catalog where you can view available foundation models, including both DigitalOcean-hosted and third-party commercial models, compare model capabilities and pricing, use routing to match inference requests to the best-fit model, and run inference using serverless or dedicated deployments.

doctl is the command-line interface for the DigitalOcean API. It supports most of the same actions available in the API and DigitalOcean Control Panel.

The Inference doctl commands are organized into the following groups:

  • Dedicated Inference: Create, list, update, and delete dedicated inference endpoints, and manage their accelerators, sizes, GPU model configuration, and access tokens. Dedicated Inference is available in public preview. You can opt in from the Feature Preview page.

  • Serverless Inference: Call models for chat completions, embeddings, images, messages, and responses, start and retrieve asynchronous invocations, list available models and regions, and manage OpenAI API keys.

  • Knowledge Base: Create, list, update, and delete agent knowledge bases, attach and detach them from agents, manage their data sources, and run indexing jobs.

Batch Inference has no doctl commands. Use the Batch Inference API instead.

For the full command tree, including groups outside Inference, see the doctl reference. Run any command with --help for its options and flags.

To install and authenticate doctl, see How to Install and Configure doctl.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.