GPU Worker Nodes

Last verified 9 Sep 2026

DigitalOcean Kubernetes (DOKS) is a Kubernetes service with a fully managed control plane, high availability, and autoscaling. DOKS integrates with standard Kubernetes toolchains and DigitalOcean’s load balancers, volumes, CPU and GPU Droplets, API, and CLI.

GPU worker nodes are built on GPU Droplets, which are powered by AMD and NVIDIA GPUs. Using GPU worker nodes in your cluster, you can:

  • Experiment and develop AI/ML applications in containerized environments
  • Run distributed AI workloads on Kubernetes
  • Scale AI inference services

Available GPU Node Pools

We offer the following GPU options for creating node pools:

AMD GPU Slug
Instinct MI300X gpu-mi300x1-192gb
Instinct MI300X (8x) gpu-mi300x8-1536gb
Instinct MI300X (8x) for multi-node setup By contract
Instinct MI325X By contract
Instinct MI325X (8x) By contract
Instinct MI325X (8x) for multi-node setup By contract
Instinct MI350X By contract
Instinct MI350X (8x) By contract
Instinct MI350X (8x) for multi-node setup By contract
NVIDIA GPU Slug
B300 By contract
B300 for multi-node setup By contract
H100 gpu-h100x1-80gb
H100 (8x) gpu-h100x8-640gb
H100 (8x) for multi-node setup By contract
H200 (8x) for multi-node setup By contract
L40s gpu-l40sx1-48gb
RTX 4000 gpu-4000adax1-20gb
RTX 6000 gpu-6000adax1-48gb
Note

For multiple 8 GPU H100 worker nodes, we support high-speed networking between GPUs from different nodes. High-speed communication uses 8x Mellanox 400GbE interfaces. Multi-node H100 setups are available by contract only.

You can use the 8-GPU configuration GPUs in a multi-node set up, where the GPUs are connected via a dedicated high-speed networking fabric. To learn how to configure high-speed networking for multi-node GPUs, see How to Use Multi-Node GPUs.

Spot GPU Node Pools public

Spot GPU node pools run worker nodes on DigitalOcean’s interruptible Spot GPU capacity, which uses available idle GPU capacity at a lower, variable rate than on-demand GPU nodes. They are best for fault-tolerant workloads that can checkpoint and resume, such as batch model training, hyperparameter tuning, and asynchronous inference. For a comparison of the two capacity tiers, see Spot GPU Droplets vs On-Demand GPU Droplets.

The following GPU shapes are available as Spot node pools:

GPU Slug
AMD Instinct MI350X gpu-mi350x1-288gb-spot
AMD Instinct MI350X (8x) gpu-mi350x8-2304gb-spot
AMD Instinct MI355X gpu-mi355x1-288gb-spot
AMD Instinct MI355X (8x) gpu-mi355x8-2304gb-spot
NVIDIA B300 (air-cooled) gpu-b300x1-288gb-spot
NVIDIA B300 (air-cooled, 8x) gpu-b300x8-2304gb-spot
NVIDIA B300 (liquid-cooled) gpu-b300x1-288gb-lc-spot
NVIDIA B300 (liquid-cooled, 8x) gpu-b300x8-2304gb-lc-spot

Spot GPU node pools are supported in the ATL1, RIC1, and MKC1 datacenter regions, with MEM1 as early access. A supported region does not guarantee current capacity, which varies by GPU model and datacenter. Spot GPU node pools draw from the same capacity pool as Spot GPU Droplets, so you can check current capacity by datacenter on the Spot GPU Droplet Capacity page.

Keep the following in mind when you run Spot GPU node pools:

  • Interruptible capacity. DigitalOcean can reclaim a Spot GPU node pool when it needs the capacity back. We aim to give at least two hours’ notice on a best-effort basis, sent by email, but less or no notice is possible in emergencies. Reclamation applies to the entire node pool, not to individual nodes.
  • Automated draining. When a reclaim signal fires, DOKS cordons and drains the pool’s nodes, respecting PodDisruptionBudgets, and emits native Kubernetes events and node conditions so your scheduler or custom hooks can react.
  • Locked pricing. The Spot rate is locked when you create the node pool and stays the same for the life of the pool, including nodes added later by autoscaling. See Kubernetes pricing.
  • Scale-down-only resizing. After you create a Spot GPU node pool, you can maintain or reduce its node count but not increase it. See Spot GPU node pool autoscaling.
  • Labels. Spot GPU nodes carry a doks.digitalocean.com/capacity-type=spot label so you can target them with node selectors or node affinity. DOKS does not apply a taint to Spot GPU nodes by default. To keep other workloads off them, add a taint to the node pool and matching tolerations to your Spot workloads.

A cluster with a GPU node pool must also have at least one CPU node pool. Because Spot capacity can be reclaimed at any time, we recommend adding an on-demand GPU node pool as a fallback. Failover is not automatic: you configure it using labels, taints, tolerations, node affinity, and the cluster autoscaler priority expander so that pods from a reclaimed Spot pool can reschedule onto the on-demand pool.

Runtime, Drivers, and Plugins

To run GPU workloads, you do not need to specify a Runtime Class.

DigitalOcean also installs and manages the required drivers to enable the GPU worker nodes as described below. For the latest versions installed, see the DOKS changelog.

AMD GPUs

For AMD GPUs, DigitalOcean installs the following on worker nodes:

Starting with DOKS 1.36.0-do.0, AMD ROCm is not installed on AMD GPU worker nodes. Use container images that bundle ROCm when your workload needs it. For AMD GPU worker images on earlier DOKS minor versions, ROCm was installed on the node; see the DOKS changelog for the ROCm versions that applied to each release.

We also recommend:

  • ROCm Device Plugin for Kubernetes for GPU discovery, health checks, configuration of GPU-enabled containers, and time slicing.

    We automatically deploy this component when you create or update a cluster. You can turn this option off by setting amd_gpu_device_plugin to false in the request body when creating or updating a cluster using the API. For an opt-in DRA alternative to the device plugin, see Dynamic Resource Allocation.

  • AMD Device Metrics Exporter for ingesting GPU metrics into your monitoring system.

    You can install this plugin by setting amd_gpu_device_metrics_exporter_plugin to true in the request body when creating or updating a cluster using the API. The plugin is installed in the kube-system namespace of the Kubernetes cluster.

Optimize Inference Model Loading on AMD GPUs

When serving models on AMD GPU workers with vLLM or SGLang, tensor shard load time can be high. Append the following flags to reduce it:

  • vLLM: Add --trust-remote-code --load-format fastsafetensors to your vllm serve command.
  • SGLang: Add --weight-loader-disable-mmap to your SGLang serve command.

NVIDIA GPUs

For NVIDIA GPUs, we install the following drivers and plugins:

We also recommend:

Additional Features

Labels and Taints

DigitalOcean applies additional labels and taints to the GPU worker nodes. For more information, see Automatic Application of Labels and Taints to Nodes.

Node Readiness Controller

DOKS also uses a node readiness controller to gate scheduling on GPU nodes until the GPU drivers are healthy. For more information, see How to Use the Node Readiness Controller in DOKS.

Scaling to Zero

You can also use the cluster autoscaler to automatically scale the GPU node pool down to zero, or use the DigitalOcean CLI or API to manually scale the node pool down to 0. Autoscaling is useful when using on-demand and for jobs like training and fine-tuning.

Dynamic Resource Allocation public

Note

Dynamic Resource Allocation for GPU node pools is in public preview.

Dynamic Resource Allocation (DRA) is an opt-in alternative to the Kubernetes device plugins for allocating GPUs to workloads. Starting with DOKS 1.36.3-do.0, DOKS can deploy DRA drivers for NVIDIA and AMD GPU node pools.

Enable a DRA driver by setting nvidia_gpu_dra_driver or amd_gpu_dra_driver to true when creating or updating a cluster using the API, doctl, or Terraform.

DRA drivers are mutually exclusive with the corresponding device plugins (nvidia_gpu_device_plugin and amd_gpu_device_plugin). You cannot enable both for the same GPU brand on a cluster.

Warning

Switching between a device plugin and a DRA driver on an existing node can leave stale kubelet state that prevents the new driver from managing GPU resources correctly. When you enable or disable a DRA driver, DOKS replaces the GPU worker nodes so each new node starts with a clean state before the selected driver is installed.

DRA remains compatible with workloads that request GPUs using device-plugin-style extended resources such as nvidia.com/gpu or amd.com/gpu. Kubernetes extended resource allocation by DRA (DRAExtendedResource) creates the corresponding ResourceClaims automatically, so you do not need to update those workloads at the same time you enable the DRA driver.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.