Running large language models (LLMs) locally on your own infrastructure provides complete data privacy, eliminates per-token API fees, and enables offline capabilities. Ollama has established itself as one of the most efficient and user-friendly runtimes for serving open-weight models such as Llama 3, Mistral, and DeepSeek.

However, running modern AI models strictly on system CPUs is notoriously slow. To achieve responsive token generation, offloading model layers to an NVIDIA GPU is crucial. While starting a single Docker container via docker run --gpus all is simple, managing production services, volume persistence, environment flags, and interconnected frontends demands a declarative Docker Compose setup.
In this comprehensive hands-on guide, you will learn step-by-step how to configure Ollama with NVIDIA GPU acceleration using Docker Compose (Compose v2) on Linux. We will walk through host driver prerequisites, proper NVIDIA Container Toolkit installation, the modern Docker Compose GPU reservation syntax, and essential troubleshooting steps for common containerized GPU pitfalls.
Technical Prerequisites
Before deploying the container stack, ensure your Linux host machine satisfies the following hardware and software requirements:
- Operating System: Ubuntu 22.04 LTS, Ubuntu 24.04 LTS, or Debian 12 / 13.
- GPU Hardware: An NVIDIA graphics card (GeForce RTX 3000/4000 series, RTX 5000/6000 series, or datacenter GPUs such as A10/A100/H100/L40S). Minimum recommended VRAM is 8 GB for 7B/8B parameter models, or 16+ GB for 14B–32B models.
- NVIDIA Proprietary Drivers: Installed on the host (version 535 or newer). Verify installation with:
nvidia-smi - Docker Engine & Docker Compose v2: Docker Engine 24.0+ and Compose v2.20+. Note that legacy
docker-compose(Python-based v1) is deprecated; you must use the moderndocker composeplugin.
Step 1: Installing the NVIDIA Container Toolkit
Docker cannot communicate with the underlying GPU hardware directly through standard namespaces. It requires the NVIDIA Container Toolkit (formerly nvidia-docker2), which exposes host CUDA drivers and runtime libraries into the container.
To configure the package repository and install the toolkit on Ubuntu or Debian, execute the following commands as root or with sudo, following the official guidelines from the NVIDIA Container Toolkit Official Installation Guide:
# 1. Configure the production repository
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
&& curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
# 2. Update package lists and install the toolkit
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
# 3. Configure the Docker daemon to use the NVIDIA runtime
sudo nvidia-ctk runtime configure --runtime=docker
# 4. Restart the Docker daemon to apply runtime changes
sudo systemctl restart docker
Step 2: Verifying GPU Acceleration Inside Docker
Before writing your Compose file, verify that Docker can successfully detect and access your GPU from within a test container:
docker run --rm --gpus all ubuntu:24.04 nvidia-smi
If everything is configured correctly, this command will output the standard nvidia-smi status table showing your GPU model, driver version, and CUDA version from inside the temporary container. If you encounter errors such as unknown flag: --gpus or could not select device driver, double-check that nvidia-ctk runtime configure was run and the Docker daemon was restarted.
Step 3: Crafting the Production docker-compose.yml
In older Docker Compose specifications (v1 / v2.3 file format), passing GPUs was often done via vendor-specific runtime: nvidia keys. In modern Docker Compose v2 specification, GPU pass-through is handled declaratively using the standard deploy.resources.reservations.devices block.
Create a dedicated directory for your Ollama stack:
mkdir -p ~/stacks/ollama && cd ~/stacks/ollama
Create a file named docker-compose.yml with the following configuration, referencing the specifications detailed in the Ollama Official Docker Documentation:
services:
ollama:
image: ollama/ollama:latest
container_name: ollama
restart: unless-stopped
ports:
- "11434:11434"
volumes:
- ./ollama_data:/root/.ollama
environment:
- OLLAMA_KEEP_ALIVE=24h
- OLLAMA_ORIGINS=*
- OLLAMA_NUM_PARALLEL=4
- OLLAMA_MAX_LOADED_MODELS=2
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
Explaining the Configuration Directives
image: ollama/ollama:latest: Pulls the official multi-architecture image. Ollama’s default Linux image already includes the required CUDA runtime libraries.volumes: - ./ollama_data:/root/.ollama: Crucial for production. Weights for modern models range from 4 GB to over 20 GB. Storing them on a persistent host volume prevents re-downloading gigabytes of data every time the container is recreated or updated.ports: - "11434:11434": Exposes Ollama’s HTTP REST API on the standard port. If you plan to expose this across a public network, place a reverse proxy like Caddy or Nginx in front with SSL and basic authentication.deploy.resources.reservations.devices:driver: nvidia: Instructs Docker to leverage the NVIDIA Container Runtime.count: all: Grants the container access to all available host GPUs.capabilities: [gpu]: Required capability flag enabling CUDA compute capabilities.
- Environment Variables:
OLLAMA_KEEP_ALIVE=24h: Keeps loaded models in GPU VRAM for 24 hours instead of unloading after the default 5-minute idle timeout.OLLAMA_ORIGINS=*: Permits Cross-Origin Resource Sharing (CORS).OLLAMA_NUM_PARALLEL=4: Enables parallel processing of up to 4 concurrent user requests.
Step 4: Starting the Ollama Service
Start the container in detached mode using Docker Compose:
docker compose up -d
Check the startup logs to ensure that Ollama detected the NVIDIA GPU and initialized the CUDA compute backend:
docker compose logs -f ollama
Step 5: Pulling and Testing Your First Model
Once the container is healthy, pull a model like llama3.1:8b:
docker compose exec ollama ollama run llama3.1:8b
While the model is generating text, open a second terminal on your host machine and monitor GPU utilization:
watch -n 1 nvidia-smi
Step 6: Verifying via the REST API
Ollama exposes a fully compatible OpenAI-like API endpoint. You can test your GPU-accelerated deployment remotely using curl:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1:8b",
"prompt": "Explain Docker Compose in two sentences.",
"stream": false
}'
Troubleshooting Common Errors
1. could not select device driver “” with capabilities: [[gpu]]
- Root Cause: The Docker daemon is not aware of the NVIDIA container runtime.
- Resolution: Re-run
sudo nvidia-ctk runtime configure --runtime=dockerand ensuresudo systemctl restart dockeris completed.
2. Ollama Runs on CPU Despite Having a GPU
- Root Cause: Insufficient free VRAM or outdated host drivers.
- Resolution: Check
docker compose logs ollamafor entries likemodel requires 22.4 GiB of VRAM but only 15.8 GiB available. Choose an appropriately quantized model (e.g.,q4_k_m).
3. Docker Compose v1 Syntax Errors
- Root Cause: Running legacy
docker-composeinstead of moderndocker compose. - Resolution: Install the
docker-compose-pluginpackage viaapt-get install docker-compose-plugin.
Conclusion & Next Steps
Deploying Ollama with NVIDIA GPU acceleration via Docker Compose creates a high-performance, reproducible foundation for all your self-hosted AI projects. By structuring your setup with persistent host volumes and the standardized Compose v2 device reservation syntax, your models remain safe across container updates and deliver peak inference speeds.
Now that your local inference engine is up and running, you can connect frontend interfaces and autonomous workflows:
- Deploy Open-WebUI: Connect a full-featured ChatGPT-like browser interface to your Ollama container over a private Docker network.
- Build an Offline RAG Stack: Integrate AnythingLLM to perform semantic search over local PDF manuals and corporate documents without any third-party cloud dependencies.
Hi, I’m Mark, the author of Clever IT Solutions: Mastering Technology for Success. I am passionate about empowering individuals to navigate the ever-changing world of information technology. With years of experience in the industry, I have honed my skills and knowledge to share with you. At Clever IT Solutions, we are dedicated to teaching you how to tackle any IT challenge, helping you stay ahead in today’s digital world. From troubleshooting common issues to mastering complex technologies, I am here to guide you every step of the way. Join me on this journey as we unlock the secrets to IT success.


