How to Deploy Ollama with Open-WebUI Pipelines and Qwen 3.6 for Advanced Autonomous Agents in Docker

AI engineer orchestrating Open-WebUI Pipelines with Ollama and Qwen 3.6 in modern workspace
Orchestrating autonomous local AI agents with Ollama, Open-WebUI Pipelines, and modern Qwen 3.6 models.

Self-hosting a local Large Language Model via Ollama provides immense privacy and eliminates per-token API billing. However, a plain chat interface quickly reaches its limits when building real-world enterprise automations. Modern applications require deterministic guardrails, multi-step agent reasoning, dynamic data retrieval (RAG), live internet search augmentation, and Python function execution. Attempting to hardcode these workflows directly into the inference engine creates brittle, unmaintainable architectures.

The standard architectural solution is Open-WebUI Pipelines. Built as an extensible, modular Python middleware framework, Pipelines sits directly between Open-WebUI and your underlying inference backends. It functions as an intelligent proxy that can intercept, modify, validate, and branch prompt-response streams in real time. Combined with Ollama running the cutting-edge Qwen 3.6 model family (renowned for top-tier reasoning, tool calling, and coding benchmarks), this containerized stack transforms a simple self-hosted chat UI into a private, autonomous agent orchestration platform. In this comprehensive guide, we configure and deploy Ollama, Open-WebUI, and the Pipelines engine using Docker Compose.

Architecture & Pipeline Interception Flow

Open-WebUI communicates with upstream LLMs using the OpenAI-compatible API standard. The Pipelines container exposes an identical OpenAI-compatible endpoint on port 9099. When Open-WebUI routes a chat request through Pipelines, the framework passes the user prompt through an ordered chain of modular Python classes: Inlet filters (preprocessing, prompt injection defense, PII scrubbing), Pipe execution (function calling, web search, agent loops against Ollama running Qwen 3.6), and Outlet filters (hallucination checks, output formatting, audit logging).

+---------------------------------------------------------------+
|                       User Browser Interface                  |
|                        Open-WebUI (Port 8080)                 |
+-------------------------------+-------------------------------+
                                |
                                | HTTP (OpenAI API Format)
                                v
+---------------------------------------------------------------+
|                 Open-WebUI Pipelines (Port 9099)              |
|  +---------------------------------------------------------+  |
|  | Inlet Filters:       PII Redaction, Prompt Sanitization |  |
|  | Pipe Logic:          RAG, Function Tools, Agent Loops   |  |
|  | Outlet Filters:      Toxicity & Hallucination Check     |  |
|  +---------------------------------------------------------+  |
+-------------------------------+-------------------------------+
                                |
                                | HTTP (Port 11434)
                                v
+---------------------------------------------------------------+
|                     Ollama Inference Engine                   |
|  - NVIDIA GPU Acceleration via Container Toolkit              |
|  - Loaded Model: Qwen 3.6 (Instruct / Coder / Vision)         |
|  - High-throughput GGUF quantization (Q4_K_M / Q8_0)          |
+---------------------------------------------------------------+

Host Prerequisites & NVIDIA GPU Support

This deployment is optimized for Linux servers (Ubuntu 24.04/22.04 LTS or Debian 12) equipped with Docker Engine (v26.0+) and Docker Compose (v2.27+). While the stack will fall back to CPU execution, an NVIDIA GPU with at least 12GB to 16GB of VRAM (e.g., RTX 3060/4060 Ti, RTX 4080, or enterprise A4000/A5000) is strongly recommended for responsive token generation with Qwen 3.6.

Verify that the NVIDIA Container Toolkit is installed and operational on your host:

# Verify GPU visibility inside Docker
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

Directory Structure & Pipeline Volume Setup

Create a dedicated workspace to store model weights, user databases, and custom Python pipeline modules:

sudo mkdir -p /opt/ai-stack/{ollama,open-webui,pipelines/pipelines}
cd /opt/ai-stack
sudo chmod -R 755 /opt/ai-stack

Production Docker Compose Specification

Create the docker-compose.yml file in /opt/ai-stack/docker-compose.yml. We link ollama, pipelines, and open-webui across a shared internal bridge network:

# /opt/ai-stack/docker-compose.yml
services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama-engine
    restart: unless-stopped
    volumes:
      - ./ollama:/root/.ollama
    ports:
      - "11434:11434"
    networks:
      - ai-network
    environment:
      - OLLAMA_KEEP_ALIVE=24h
      - OLLAMA_NUM_PARALLEL=4
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

  pipelines:
    image: ghcr.io/open-webui/pipelines:main
    container_name: open-webui-pipelines
    restart: unless-stopped
    volumes:
      - ./pipelines/pipelines:/app/pipelines
    ports:
      - "9099:9099"
    networks:
      - ai-network
    environment:
      - PIPELINES_URLS=https://github.com/open-webui/pipelines/archive/refs/heads/main.zip
      - PIPELINES_REQUIREMENTS_PATH=/app/pipelines/requirements.txt
    depends_on:
      - ollama
    deploy:
      resources:
        limits:
          memory: 2048M
          cpus: "2.0"

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui-frontend
    restart: unless-stopped
    volumes:
      - ./open-webui:/app/backend/data
    ports:
      - "8080:8080"
    networks:
      - ai-network
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - OPENAI_API_BASE_URL=http://pipelines:9099
      - OPENAI_API_KEY=0p3n-w3bu1-p1p3l1n3s
      - WEBUI_AUTH=true
      - ENABLE_SIGNUP=false
    depends_on:
      - pipelines
    deploy:
      resources:
        limits:
          memory: 2048M
          cpus: "2.0"

networks:
  ai-network:
    name: ai-network
    driver: bridge

Pulling and Verifying the Qwen 3.6 Model

Start the Docker Compose services in detached mode:

cd /opt/ai-stack
docker compose up -d

# Verify all three containers are actively running
docker compose ps

Next, pull the latest Qwen 3.6 model into Ollama. The 14B parameter variant strikes an optimal balance between reasoning accuracy and high token throughput on standard 16GB VRAM GPUs, while the 7B or 32B variants can be chosen depending on hardware capacity:

# Pull the modern Qwen 3.6 reasoning model
docker exec -it ollama-engine ollama pull qwen3.6:14b

# Verify the model is downloaded and registered
docker exec -it ollama-engine ollama list

Authoring a Custom Autonomous Pipeline Module

The core strength of Open-WebUI Pipelines is its plug-and-play Python architecture. Any script placed into /opt/ai-stack/pipelines/pipelines/ that implements the Pipeline class is automatically registered as a distinct, callable model or filter inside Open-WebUI.

Below is a production-ready custom pipeline that implements an automated System Prompt Injection & Rate-Limiting Guardrail for Qwen 3.6. Create /opt/ai-stack/pipelines/pipelines/guardrail_qwen.py:

# /opt/ai-stack/pipelines/pipelines/guardrail_qwen.py
import os
import requests
from typing import List, Union, Generator, Iterator
from pydantic import BaseModel

class Pipeline:
    class Valves(BaseModel):
        OLLAMA_HOST: str = "http://ollama:11434"
        BASE_MODEL: str = "qwen3.6:14b"
        ENFORCE_STRICT_JSON: bool = False

    def __init__(self):
        self.type = "pipe"
        self.id = "qwen_autonomous_guardrail"
        self.name = "Qwen 3.6 Autonomous Guardrail"
        self.valves = self.Valves()

    async def on_startup(self):
        print(f"[{self.name}] Pipeline initialized and connected to {self.valves.OLLAMA_HOST}")

    async def on_shutdown(self):
        print(f"[{self.name}] Pipeline shutting down.")

    def pipe(
        self, user_message: str, model_id: str, messages: List[dict], body: dict
    ) -> Union[str, Generator, Iterator]:
        # Inject standard security preamble for enterprise agent behavior
        system_guard = {
            "role": "system",
            "content": (
                "You are an autonomous engineering assistant powered by Qwen 3.6. "
                "Always adhere to strict factual accuracy. Provide production-ready, "
                "hardened code snippets with explanatory commentary. Avoid pleasantries."
            )
        }

        # Prepend system guardrail if not already present
        if not any(m.get("role") == "system" for m in messages):
            messages.insert(0, system_guard)

        payload = {
            "model": self.valves.BASE_MODEL,
            "messages": messages,
            "stream": body.get("stream", False)
        }

        response = requests.post(
            f"{self.valves.OLLAMA_HOST}/api/chat",
            json=payload,
            stream=body.get("stream", False),
            timeout=180
        )

        if response.status_code != 200:
            return f"Error communicating with Ollama: HTTP {response.status_code} - {response.text}"

        if body.get("stream", False):
            def stream_generator():
                for line in response.iter_lines():
                    if line:
                        import json
                        chunk = json.loads(line)
                        yield chunk.get("message", {}).get("content", "")
            return stream_generator()
        else:
            data = response.json()
            return data.get("message", {}).get("content", "")

Restart the Pipelines service to compile and load the new module:

docker compose restart pipelines
docker compose logs -f pipelines | grep "Pipeline initialized"

Connecting Pipelines Inside Open-WebUI

Navigate to http://<your-server-ip>:8080 in your browser. Register your primary administrator account.

  • Go to Admin Panel -> Settings -> Connections.
  • Under OpenAI API, ensure the endpoint points to http://pipelines:9099 with the API key 0p3n-w3bu1-p1p3l1n3s.
  • Click the verify/refresh button. Open-WebUI will immediately query Pipelines and discover your newly defined Qwen 3.6 Autonomous Guardrail model.
  • In the main chat window, select Qwen 3.6 Autonomous Guardrail from the top model selector and test a multi-turn conversation.

Production Hardening & Best Practices

  • Keep Inference & WebUI Internal: Never expose port 11434 (Ollama) or port 9099 (Pipelines) directly to the public internet. Ollama has no built-in authentication mechanism by default. Keep all three containers attached exclusively to the internal Docker bridge network, exposing only Open-WebUI (port 8080) behind a hardened TLS reverse proxy (Caddy or Cloudflare Tunnels).
  • Configure Memory & VRAM Offloading: The OLLAMA_KEEP_ALIVE=24h environment variable keeps Qwen 3.6 loaded inside GPU VRAM, eliminating the 10-second cold-load delay on the first user prompt of the day. If your host has limited RAM, reduce this setting to OLLAMA_KEEP_ALIVE=15m.
  • Sanitize Python Dependencies: When utilizing third-party community pipelines (such as LangChain RAG or DuckDuckGo search plugins), always review the required dependencies in requirements.txt before restarting the Pipelines container to prevent unauthorized arbitrary code execution.
  • Disable Open Registration: Prevent unauthorized local token consumption by ensuring ENABLE_SIGNUP=false is set in your docker-compose.yml file after creating your administrator account.

Troubleshooting Common Failures

1. Ollama Falls Back to CPU (Extremely Slow Generation)

Root Cause: The NVIDIA Container Toolkit is not registered with Docker’s daemon runtime, or the GPU reservations block is missing in docker-compose.yml.
Solution: Verify that /etc/docker/daemon.json includes the nvidia runtime. Restart the Docker daemon via sudo systemctl restart docker. Check Ollama GPU utilization in logs: docker compose logs ollama | grep -i "gpu".

2. Open-WebUI Reports “Cannot Connect to OpenAI API”

Root Cause: Open-WebUI cannot resolve the pipelines container hostname, or the API key does not match.
Solution: Confirm both containers reside on the same Docker network (ai-network). Ensure OPENAI_API_BASE_URL=http://pipelines:9099 uses the container name, not localhost, as localhost refers to the Open-WebUI container itself.

3. Custom Pipeline Fails to Load with ModuleNotFoundError

Root Cause: Your custom Python script requires external packages that are not bundled into the base Pipelines image.
Solution: Create a requirements.txt file in /opt/ai-stack/pipelines/pipelines/requirements.txt listing your dependencies (e.g., beautifulsoup4, duckduckgo_search). When the container boots, it automatically installs these packages before initializing the pipeline modules.

Conclusion

By pairing Ollama and Qwen 3.6 with Open-WebUI Pipelines, you elevate local AI from a simple interactive chatbot into a programmable, modular workflow ecosystem. With Python-based inlet/outlet filters, deterministic system prompt injection, and seamless agent tool routing, you gain absolute control over model inputs and outputs—all hosted entirely on your private hardware with zero vendor telemetry or recurring API expenses.