Self-Host a Private ChatGPT: Open-WebUI + Ollama via Docker Compose

Commercial cloud AI chatbots like ChatGPT, Claude, and Gemini have transformed productivity, but they present significant data privacy risks and ongoing subscription costs for technical teams and privacy-conscious users. When you submit proprietary code, sensitive business spreadsheets, or personal notes to cloud services, your data is processed on remote infrastructure outside your perimeter.

Software engineer testing self-hosted Open-WebUI and Ollama chat interface
Self-Host a Private ChatGPT: Open-WebUI + Ollama via Docker Compose 3

The open-source community has developed a compelling alternative: combining Ollama (the high-performance local model inference server) with Open-WebUI (formerly Ollama WebUI), a feature-rich, responsive web interface that replicates and often surpasses the ChatGPT user experience.

In this tutorial, you will learn how to orchestrate a production-ready, fully private AI chat platform on your own server or homelab using Docker Compose. We will configure seamless container networking, persistent storage for models and conversation history, role-based user management, and optional GPU passthrough.

Why Open-WebUI + Ollama is the Premier Self-Hosted Stack

Running Ollama solely via the command line is great for quick terminal experiments, but lacks the collaborative interface modern users expect. Open-WebUI bridges this gap by providing:

  • Familiar Chat Experience: Full Markdown rendering, syntax highlighting for code blocks, LaTeX math support, and streaming responses.
  • Granular Model Management: Download, update, switch, and delete models directly from the web interface without touching terminal commands.
  • Document Ingestion & RAG: Built-in semantic search over uploaded PDFs, Markdown files, and text documents.
  • Multi-User Collaboration: Role-Based Access Control (Admin vs. User), user invitation systems, and isolated chat histories.
  • Extensibility: Support for web search engines (SearXNG, Google, DuckDuckGo), custom system prompts, and multi-modal models (vision models like LLaVA).

If you have an NVIDIA GPU, make sure you review our previous deep dive on running Ollama with NVIDIA GPU acceleration in Docker Compose to maximize your tokens per second before launching this multi-container stack.

Technical Prerequisites

Before deploying the stack, ensure you have:

  1. Host Environment: A Linux server or desktop running Ubuntu 22.04 / 24.04 LTS or Debian 12 / 13.
  2. Docker Engine & Docker Compose v2: Docker Engine 24.0+ with the docker compose CLI plugin installed.
  3. Hardware Sizing:
    • CPU-only: Minimum 4 vCPUs and 16 GB of RAM (suitable for small 3B–8B quantized models).
    • GPU-accelerated (Recommended): Dedicated NVIDIA GPU (RTX 3060 12GB, RTX 4070/4080, or enterprise cards) with the NVIDIA Container Toolkit installed.
  4. Storage: At least 30–50 GB of free NVMe or SSD disk space for storing model weights and chat databases.

Step 1: Planning Container Networking & Storage

When orchestrating Open-WebUI alongside Ollama in Docker, container communication is the most frequent stumbling block.

In single-container setups, tutorials often tell users to set OLLAMA_BASE_URL=http://host.docker.internal:11434. However, host.docker.internal requires special extra_hosts mappings on Linux. The cleanest, most secure, and production-grade approach is to place both containers onto a user-defined Docker bridge network. Within this isolated network, Open-WebUI can reach Ollama directly via its container hostname: http://ollama:11434.

We also must define two separate persistent host volumes:

  1. ollama_data: Houses downloaded GGUF model weights in /root/.ollama.
  2. openwebui_data: Houses Open-WebUI’s internal SQLite database, user accounts, RAG vector embeddings, and session state in /app/backend/data.

Step 2: Creating the Project Directory and Compose File

Create a clean directory for your private chat stack:

mkdir -p ~/stacks/private-chat && cd ~/stacks/private-chat

Create a file named docker-compose.yml and paste the following verified configuration, which follows the architecture laid out in the Open-WebUI Official Documentation and the Ollama Documentation:

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "127.0.0.1:11434:11434"
    volumes:
      - ./ollama_data:/root/.ollama
    environment:
      - OLLAMA_KEEP_ALIVE=24h
      - OLLAMA_NUM_PARALLEL=4
    networks:
      - ai-network
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    restart: unless-stopped
    ports:
      - "3000:8080"
    volumes:
      - ./openwebui_data:/app/backend/data
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - WEBUI_SECRET_KEY=change_this_to_a_random_secure_hex_key
      - DEFAULT_MODELS=llama3.1:8b
      - ENABLE_SIGNUP=true
    depends_on:
      - ollama
    networks:
      - ai-network

networks:
  ai-network:
    driver: bridge

Step 3: Understanding Crucial Configuration Directives

Let’s dissect the essential parameters in this configuration:

1. Network Binding & Security

  • ports: - "127.0.0.1:11434:11434" for Ollama: We bind Ollama strictly to 127.0.0.1 on the host. This prevents unauthenticated users on your local network from directly querying Ollama’s raw REST API, forcing all user traffic through Open-WebUI’s authentication layer.
  • ports: - "3000:8080" for Open-WebUI: Maps the container’s internal port 8080 to host port 3000. You will access your chat interface at http://your-server-ip:3000.

2. Environment Variables

  • OLLAMA_BASE_URL=http://ollama:11434: Informs Open-WebUI where to find the inference backend. Because both containers share ai-network, Docker’s internal DNS resolves ollama directly to the correct container IP.
  • WEBUI_SECRET_KEY: A persistent cryptographic key used to sign session cookies and JWT tokens. Replace the placeholder with a secure 32-character hex string generated via openssl rand -hex 32. If omitted, Open-WebUI regenerates a random key on every container restart, which logs all active users out.
  • ENABLE_SIGNUP=true: When deploying for the first time, keep this set to true so you can register your initial admin account. After creating your account, you can change this to false to block unauthorized public signups.

Step 4: Launching the Stack

Generate your secret key and launch the services in background mode:

# Generate a secret key
export RANDOM_SECRET=$(openssl rand -hex 32)
sed -i "s/change_this_to_a_random_secure_hex_key/$RANDOM_SECRET/" docker-compose.yml

# Start the stack
docker compose up -d

Monitor container initialization with docker compose logs:

docker compose logs -f

You should observe Ollama initializing its listening socket on port 11434, followed by Open-WebUI completing database migrations and launching its ASGI server on port 8080.

Step 5: Initial Setup & Admin Account Creation

  1. Open your web browser and navigate to:
    http://localhost:3000
    (Or replace localhost with your server’s LAN IP address).
  2. You will be greeted by the Open-WebUI welcome screen. Click Sign Up.
  3. Important: The first user account registered automatically receives Admin privileges. Enter your name, email, and a strong master password.
  4. Once logged in, navigate to Admin Panel > Settings > General:
    • If this instance is only for yourself or a closed team, toggle Enable New Signups to OFF.
    • You can manage pending user access or generate invite links directly from this dashboard.

Step 6: Pulling Models Directly from Open-WebUI

One of Open-WebUI’s standout features is that you do not need terminal access to install new LLMs.

  1. In the top-right corner, click on your profile picture and open Admin Panel > Settings > Connections.
  2. Verify that the Ollama connection status indicates Connected to http://ollama:11434.
  3. Go to Settings > Models or click on the model selector at the top of the main chat window.
  4. In the Pull a model from Ollama.com field, enter a model tag:
    • llama3.1:8b (Meta’s versatile 8-billion parameter model)
    • deepseek-r1:8b (Specialized reasoning model)
    • qwen2.5-coder:7b (High-performance code generation model)
  5. Click the Download (Pull) icon. Open-WebUI will stream download progress in real-time. The weights are automatically persisted in ./ollama_data/models.

Step 7: Testing Chat, Code Execution, and Vision

Once your model has finished downloading, click New Chat:

  1. Select your downloaded model from the dropdown.
  2. Send a query with Markdown formatting:
    Write a Python script that calculates prime numbers using the Sieve of Eratosthenes, and explain the time complexity.
  3. Observe the response: Open-WebUI displays formatted code with a one-click copy button, line numbers, and instant token generation speed.

If you download a multi-modal model like llava:7b or minicpm-v:latest, you can drag and drop images directly into the chat prompt to ask visual questions, extract text from receipts, or analyze system architecture diagrams.

Advanced Feature: Document Q&A (Local RAG)

Open-WebUI features an integrated Retrieval-Augmented Generation (RAG) engine powered by ChromaDB.

  1. In any chat window, click the + (Attach Document) paperclip icon or drag in a PDF / text file.
  2. Open-WebUI automatically extracts text chunks, computes embeddings using Ollama’s local embedding pipeline (or a fast sentence-transformer model), and saves them in ./openwebui_data.
  3. Ask specific questions about the document:
    What are the server maintenance SLA terms outlined in section 4 of the attached contract?

The model will cite exact passages from your document while answering, without sending a single byte of your data to an external API.

Troubleshooting Common Issues

1. “WebUI: Server Connection Error” or Cannot Reach Ollama

  • Symptom: Open-WebUI displays a red banner saying it cannot connect to Ollama.
  • Fix: Ensure both containers are on ai-network. From your host, run docker compose exec open-webui curl -s http://ollama:11434/api/tags. If this command fails or times out, verify that the service name in docker-compose.yml matches OLLAMA_BASE_URL=http://ollama:11434.

2. Slow Response Times or Freezing

  • Symptom: Responses take minutes to generate or stall mid-sentence.
  • Fix: If running on CPU, ensure you are using smaller quantized models (like 3B or 7B Q4_K_M). If running on an NVIDIA GPU, verify that the deploy.resources.reservations.devices block is active and that running watch nvidia-smi on the host displays active VRAM allocation under ollama_llama_server.

3. Logins Expire on Every Server Reboot

  • Symptom: Users are repeatedly logged out whenever Docker restarts.
  • Fix: You left WEBUI_SECRET_KEY blank or unset. Specify a permanent 64-character hex key in docker-compose.yml so session hashes remain consistent across container lifecycle events.

Conclusion & Next Steps

By combining Open-WebUI and Ollama within a declarative Docker Compose structure, you now possess a private, self-contained AI platform that rivals commercial cloud services in both capability and aesthetics. All models, conversation histories, and embedded documents remain strictly under your control.

In our next DIY guide, we will explore how to take local document intelligence even further by constructing a specialized, multi-workspace RAG pipeline using AnythingLLM alongside local embedding vectors.