Deploy LiteLLM Proxy in Docker: Load-Balancing, Fallbacks & API Keys for Ollama

Deploy LiteLLM Proxy with Docker Compose to create a unified, OpenAI-compatible API gateway for local Ollama instances. Configure load-balancing, automatic fallbacks, rate limiting, and virtual API keys.

A system administrator configuring LiteLLM proxy for load balancing and API keys with Ollama in Docker Compose
Deploy LiteLLM Proxy in Docker: Load-Balancing, Fallbacks & API Keys for Ollama 3

Running local AI models with Ollama has revolutionized private inference. However, as organizations and engineering teams expand their usage, direct point-to-point connections between client applications and individual Ollama servers quickly hit operational bottlenecks. A single Ollama server can become saturated under concurrent prompt loads, lacks granular API key tracking, does not natively balance requests across multiple GPU worker nodes, and cannot automatically fall back to cloud providers when local compute is overwhelmed.

LiteLLM Proxy provides the missing enterprise gateway layer. Acting as an ultra-fast, OpenAI-compatible proxy, LiteLLM sits between your applications and multiple LLM backends. In this comprehensive guide, we will set up LiteLLM Proxy in Docker Compose, configure multi-node load balancing across local Ollama instances, implement automatic failover routing, and issue managed virtual API keys with usage quotas.

Why Use LiteLLM Proxy in Front of Ollama?

LiteLLM Proxy translates requests into the universal OpenAI API schema (/v1/chat/completions, /v1/embeddings, /v1/models). Integrating it into your homelab or private enterprise infrastructure delivers several core benefits:

  • Load Balancing Across GPU Nodes: Distribute concurrent inference requests across multiple Ollama servers using round-robin, least-busy, or latency-based routing.
  • Automatic Failovers & Cooldowns: If an Ollama node runs out of VRAM (CUDA OOM error) or loses network connectivity, LiteLLM automatically retries the request against a secondary node or fallback model.
  • Virtual API Key Management: Create scoped API keys with rate limits (RPM/TPM), budget caps (max spend/tokens), and model access whitelists.
  • Drop-in Application Compatibility: Any software built for OpenAI (LibreChat, Open-WebUI, Cursor, LangChain, AutoGen) connects to your self-hosted models by simply updating the OPENAI_BASE_URL.
  • Detailed Audit & Telemetry Logging: Track request latencies, token consumption, and errors with Prometheus metrics, OpenTelemetry, or PostgreSQL storage.

Architecture: The LiteLLM Gateway Pattern

The diagram below demonstrates how client traffic flows through LiteLLM Proxy to distributed Ollama inference instances:

+--------------------------------------------------------------------+
|                         Client Layer                               |
|       (Open-WebUI / Coding Agents / Internal Microservices)        |
+--------------------------------------------------------------------+
                                   |
                                   | HTTP / HTTPS (Bearer Token Auth)
                                   v
+--------------------------------------------------------------------+
|                     LiteLLM Proxy (Port 4000)                      |
|                                                                    |
|  - Virtual Key Validation & Rate Limiting                         |
|  - Load Balancer & Health Checks (Cooldown Tracking)               |
|  - Unified OpenAI Schema Translator                                |
+--------------------------------------------------------------------+
              |                                        |
              | Internal Docker Bridge / LAN           |
              v                                        v
+-----------------------------+          +---------------------------+
|    Ollama Node A (Local)    |          |   Ollama Node B (Remote)  |
|  http://ollama-a:11434      |          |  http://192.168.1.50:11434|
|  (NVIDIA RTX 4090 - GPU 0)  |          |  (NVIDIA RTX 3090 - GPU 1)|
+-----------------------------+          +---------------------------+

Step 1: Preparing Directory Structure and Database

LiteLLM Proxy utilizes PostgreSQL to persist user keys, spend tracking, and audit logs. Create a dedicated project directory:

mkdir -p ~/litellm-docker/config ~/litellm-docker/pgdata
cd ~/litellm-docker

Create an environment file .env with secure secrets:

cat <<EOF > .env
LITELLM_MASTER_KEY=sk-$(openssl rand -hex 16)
POSTGRES_USER=litellm
POSTGRES_PASSWORD=$(openssl rand -hex 16)
POSTGRES_DB=litellm
EOF

Step 2: Defining the LiteLLM Configuration (config.yaml)

The core routing logic is defined in config/config.yaml. Here, we register a virtual model identifier (llama-3-fleet) backed by two distinct Ollama instances with weighted load balancing and automatic cooldowns:

model_list:
  # Primary Model Group with Load Balancing
  - model_name: llama-3-fleet
    litellm_params:
      model: ollama/llama3.1:8b
      api_base: http://ollama:11434
      rpm: 60
    model_info:
      id: "node-local"

  - model_name: llama-3-fleet
    litellm_params:
      model: ollama/llama3.1:8b
      # Example of a secondary server on your local LAN
      api_base: http://192.168.1.50:11434
      rpm: 60
    model_info:
      id: "node-worker"

  # Fallback Model: Smaller model if primary fleet is overloaded
  - model_name: fallback-fast
    litellm_params:
      model: ollama/qwen2.5:3b
      api_base: http://ollama:11434

router_settings:
  routing_strategy: "least-busy" # Options: simple-shuffle, least-busy, usage-based-routing
  model_group_alias:
    "gpt-4o-mini": "llama-3-fleet" # Transparently redirect cloud requests to local hardware
  cooldown_time: 30 # Seconds to wait before retrying an unhealthy endpoint
  num_retries: 3
  timeout: 120

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL

Step 3: Docker Compose Deployment

Now, create docker-compose.yml to run LiteLLM Proxy, PostgreSQL, and a local Ollama container in an integrated stack:

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "127.0.0.1:11434:11434"
    volumes:
      - ollama_models:/root/.ollama
    # GPU acceleration (remove deploy block if running CPU-only)
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    networks:
      - gateway_net

  postgres:
    image: postgres:16-alpine
    container_name: litellm-postgres
    restart: unless-stopped
    environment:
      POSTGRES_USER: ${POSTGRES_USER}
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
      POSTGRES_DB: ${POSTGRES_DB}
    volumes:
      - ./pgdata:/var/lib/postgresql/data
    networks:
      - gateway_net

  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    container_name: litellm-proxy
    restart: unless-stopped
    depends_on:
      - postgres
      - ollama
    ports:
      - "0.0.0.0:4000:4000"
    environment:
      - LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
      - DATABASE_URL=postgresql://${POSTGRES_USER}:${POSTGRES_PASSWORD}@postgres:5432/${POSTGRES_DB}
      - STORE_MODEL_IN_DB=True
    volumes:
      - ./config/config.yaml:/app/config.yaml
    command: ["--config", "/app/config.yaml", "--port", "4000", "--num_workers", "4"]
    networks:
      - gateway_net

networks:
  gateway_net:
    name: gateway_net
    driver: bridge

volumes:
  ollama_models:
    name: ollama_models

Step 4: Launching and Testing the Proxy

Start the services:

docker compose up -d

Pull the required model in Ollama:

docker exec -it ollama ollama pull llama3.1:8b
docker exec -it ollama ollama pull qwen2.5:3b

Retrieve your Master Key from the .env file:

grep LITELLM_MASTER_KEY .env

Send a standard OpenAI-compatible test request to port 4000:

curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_LITELLM_MASTER_KEY" \
  -d '{
    "model": "llama-3-fleet",
    "messages": [
      {"role": "user", "content": "Explain what a reverse proxy does in 2 sentences."}
    ],
    "temperature": 0.7
  }'

Step 5: Generating Managed Virtual Keys

Instead of distributing your master admin key, generate individual virtual API keys with defined token limits and model restrictions via LiteLLM’s management API:

curl -X POST http://localhost:4000/key/generate \
  -H "Authorization: Bearer YOUR_LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "models": ["llama-3-fleet"],
    "duration": "30d",
    "max_budget": 0.0,
    "rpm_limit": 30,
    "metadata": {"team": "frontend-devs"}
  }'

The response returns a scoped key (starting with sk-...) that you can safely distribute to team members or third-party web apps.

Production Hardening Checklist

  1. Reverse Proxy with TLS: Place Caddy or Nginx in front of port 4000 to enforce HTTPS and prevent plain-text bearer token transmission over the network.
  2. Set Worker Concurrency: Tune the --num_workers flag based on your CPU core count to handle high-concurrency client polling.
  3. Monitor Ollama VRAM: Ollama will queue incoming requests if VRAM is fully allocated. Use least-busy routing in LiteLLM so new requests automatically route to underutilized GPU workers.

Conclusion

Deploying LiteLLM Proxy transforms isolated local Ollama instances into an elastic, resilient, and manageable AI gateway. You achieve the operational control of public cloud AI APIs—including load balancing, telemetry, and scoped authentication—while retaining the privacy and cost-efficiency of self-hosted open-source models.