How to Run Ollama with NVIDIA GPU Acceleration in Docker Compose

Running large language models (LLMs) locally on your own infrastructure provides complete data privacy, eliminates per-token API fees, and enables offline capabilities. Ollama has established itself as one of the most efficient and user-friendly runtimes for serving open-weight models such as Llama 3, Mistral, and DeepSeek.

IT engineer configuring Ollama with NVIDIA GPU acceleration in Docker Compose
How to Run Ollama with NVIDIA GPU Acceleration in Docker Compose 3

However, running modern AI models strictly on system CPUs is notoriously slow. To achieve responsive token generation, offloading model layers to an NVIDIA GPU is crucial. While starting a single Docker container via docker run --gpus all is simple, managing production services, volume persistence, environment flags, and interconnected frontends demands a declarative Docker Compose setup.

In this comprehensive hands-on guide, you will learn step-by-step how to configure Ollama with NVIDIA GPU acceleration using Docker Compose (Compose v2) on Linux. We will walk through host driver prerequisites, proper NVIDIA Container Toolkit installation, the modern Docker Compose GPU reservation syntax, and essential troubleshooting steps for common containerized GPU pitfalls.

Technical Prerequisites

Before deploying the container stack, ensure your Linux host machine satisfies the following hardware and software requirements:

  1. Operating System: Ubuntu 22.04 LTS, Ubuntu 24.04 LTS, or Debian 12 / 13.
  2. GPU Hardware: An NVIDIA graphics card (GeForce RTX 3000/4000 series, RTX 5000/6000 series, or datacenter GPUs such as A10/A100/H100/L40S). Minimum recommended VRAM is 8 GB for 7B/8B parameter models, or 16+ GB for 14B–32B models.
  3. NVIDIA Proprietary Drivers: Installed on the host (version 535 or newer). Verify installation with:
    nvidia-smi
  4. Docker Engine & Docker Compose v2: Docker Engine 24.0+ and Compose v2.20+. Note that legacy docker-compose (Python-based v1) is deprecated; you must use the modern docker compose plugin.

Step 1: Installing the NVIDIA Container Toolkit

Docker cannot communicate with the underlying GPU hardware directly through standard namespaces. It requires the NVIDIA Container Toolkit (formerly nvidia-docker2), which exposes host CUDA drivers and runtime libraries into the container.

To configure the package repository and install the toolkit on Ubuntu or Debian, execute the following commands as root or with sudo, following the official guidelines from the NVIDIA Container Toolkit Official Installation Guide:

# 1. Configure the production repository
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
  && curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
    sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
    sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

# 2. Update package lists and install the toolkit
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit

# 3. Configure the Docker daemon to use the NVIDIA runtime
sudo nvidia-ctk runtime configure --runtime=docker

# 4. Restart the Docker daemon to apply runtime changes
sudo systemctl restart docker

Step 2: Verifying GPU Acceleration Inside Docker

Before writing your Compose file, verify that Docker can successfully detect and access your GPU from within a test container:

docker run --rm --gpus all ubuntu:24.04 nvidia-smi

If everything is configured correctly, this command will output the standard nvidia-smi status table showing your GPU model, driver version, and CUDA version from inside the temporary container. If you encounter errors such as unknown flag: --gpus or could not select device driver, double-check that nvidia-ctk runtime configure was run and the Docker daemon was restarted.

Step 3: Crafting the Production docker-compose.yml

In older Docker Compose specifications (v1 / v2.3 file format), passing GPUs was often done via vendor-specific runtime: nvidia keys. In modern Docker Compose v2 specification, GPU pass-through is handled declaratively using the standard deploy.resources.reservations.devices block.

Create a dedicated directory for your Ollama stack:

mkdir -p ~/stacks/ollama && cd ~/stacks/ollama

Create a file named docker-compose.yml with the following configuration, referencing the specifications detailed in the Ollama Official Docker Documentation:

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ./ollama_data:/root/.ollama
    environment:
      - OLLAMA_KEEP_ALIVE=24h
      - OLLAMA_ORIGINS=*
      - OLLAMA_NUM_PARALLEL=4
      - OLLAMA_MAX_LOADED_MODELS=2
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

Explaining the Configuration Directives

  • image: ollama/ollama:latest: Pulls the official multi-architecture image. Ollama’s default Linux image already includes the required CUDA runtime libraries.
  • volumes: - ./ollama_data:/root/.ollama: Crucial for production. Weights for modern models range from 4 GB to over 20 GB. Storing them on a persistent host volume prevents re-downloading gigabytes of data every time the container is recreated or updated.
  • ports: - "11434:11434": Exposes Ollama’s HTTP REST API on the standard port. If you plan to expose this across a public network, place a reverse proxy like Caddy or Nginx in front with SSL and basic authentication.
  • deploy.resources.reservations.devices:
    • driver: nvidia: Instructs Docker to leverage the NVIDIA Container Runtime.
    • count: all: Grants the container access to all available host GPUs.
    • capabilities: [gpu]: Required capability flag enabling CUDA compute capabilities.
  • Environment Variables:
    • OLLAMA_KEEP_ALIVE=24h: Keeps loaded models in GPU VRAM for 24 hours instead of unloading after the default 5-minute idle timeout.
    • OLLAMA_ORIGINS=*: Permits Cross-Origin Resource Sharing (CORS).
    • OLLAMA_NUM_PARALLEL=4: Enables parallel processing of up to 4 concurrent user requests.

Step 4: Starting the Ollama Service

Start the container in detached mode using Docker Compose:

docker compose up -d

Check the startup logs to ensure that Ollama detected the NVIDIA GPU and initialized the CUDA compute backend:

docker compose logs -f ollama

Step 5: Pulling and Testing Your First Model

Once the container is healthy, pull a model like llama3.1:8b:

docker compose exec ollama ollama run llama3.1:8b

While the model is generating text, open a second terminal on your host machine and monitor GPU utilization:

watch -n 1 nvidia-smi

Step 6: Verifying via the REST API

Ollama exposes a fully compatible OpenAI-like API endpoint. You can test your GPU-accelerated deployment remotely using curl:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1:8b",
  "prompt": "Explain Docker Compose in two sentences.",
  "stream": false
}'

Troubleshooting Common Errors

1. could not select device driver “” with capabilities: [[gpu]]

  • Root Cause: The Docker daemon is not aware of the NVIDIA container runtime.
  • Resolution: Re-run sudo nvidia-ctk runtime configure --runtime=docker and ensure sudo systemctl restart docker is completed.

2. Ollama Runs on CPU Despite Having a GPU

  • Root Cause: Insufficient free VRAM or outdated host drivers.
  • Resolution: Check docker compose logs ollama for entries like model requires 22.4 GiB of VRAM but only 15.8 GiB available. Choose an appropriately quantized model (e.g., q4_k_m).

3. Docker Compose v1 Syntax Errors

  • Root Cause: Running legacy docker-compose instead of modern docker compose.
  • Resolution: Install the docker-compose-plugin package via apt-get install docker-compose-plugin.

Conclusion & Next Steps

Deploying Ollama with NVIDIA GPU acceleration via Docker Compose creates a high-performance, reproducible foundation for all your self-hosted AI projects. By structuring your setup with persistent host volumes and the standardized Compose v2 device reservation syntax, your models remain safe across container updates and deliver peak inference speeds.

Now that your local inference engine is up and running, you can connect frontend interfaces and autonomous workflows:

  • Deploy Open-WebUI: Connect a full-featured ChatGPT-like browser interface to your Ollama container over a private Docker network.
  • Build an Offline RAG Stack: Integrate AnythingLLM to perform semantic search over local PDF manuals and corporate documents without any third-party cloud dependencies.