How to Self-Host Continue.dev with Ollama and Qwen 3.6 in Docker Compose for Private AI Coding

Deploy a private, zero-data-leakage AI coding assistant in VS Code and JetBrains using Continue.dev, Ollama, and Qwen 3.6 in Docker Compose. Includes GPU passthrough, tab autocomplete, and local codebase RAG.

Developer using Continue.dev AI coding assistant connected to local Ollama and Qwen 3.6 running in Docker Compose
Run Continue.dev connected to an air-gapped or private Ollama container with state-of-the-art Qwen 3.6 open-weights models.

Modern software development teams increasingly rely on AI-assisted coding tools for inline tab autocomplete, multi-file refactoring, docstring generation, and architectural discussions. While SaaS-based proprietary assistants like GitHub Copilot, Cursor, and Claude Dev deliver remarkable productivity gains, they present severe privacy, compliance, and Intellectual Property (IP) dilemmas for privacy-conscious organizations, healthcare providers, financial institutions, and homelab engineers. Sending proprietary source trees, API keys, database connection strings, and unreleased business logic to third-party cloud infrastructure is often strictly prohibited by corporate governance and strict data privacy regulations.

Fortunately, open-weights coding models have made a monumental leap. The modern Qwen 3.6 model family (incorporating dedicated coder architectures) rivals top-tier proprietary APIs in code synthesis, function calling, and deep contextual reasoning across Python, TypeScript, Rust, Go, and C++. When paired with Ollama running in Docker Compose and Continue.dev—the leading open-source IDE extension for VS Code and JetBrains—you can deploy a completely private, zero-data-leakage AI coding environment that runs locally on your workstation or self-hosted GPU homelab server.

In this guide, you will learn how to architect, deploy, and harden a production-grade self-hosted Continue.dev setup powered by Ollama and Qwen 3.6. We cover NVIDIA GPU hardware passthrough, optimized multi-concurrency environment tuning, Continue configuration schemas for chat and tab-autocomplete, local codebase embeddings for full repository indexing, and essential troubleshooting techniques.

Architecture Overview: Private AI Pair Programming

A resilient self-hosted AI coding infrastructure separates the client IDE interface from the resource-intensive inference engine. Rather than running inference directly inside developer laptops, hosting the Ollama server on a dedicated Docker host (equipped with an NVIDIA RTX 3090, 4090, or data center GPU) allows entire engineering teams to share computational resources securely over encrypted Tailscale meshes or private local networks.

+-----------------------------------------------------------------------+
|  Developer Workstation (VS Code / Cursor / JetBrains IDE)             |
|                                                                       |
|   +---------------------------------------------------------------+   |
|   | Continue.dev Extension                                        |   |
|   |                                                               |   |
|   |  [Tab Autocomplete]  --> FIM Queries (Qwen 3.6 Coder 7B)      |   |
|   |  [Interactive Chat]  --> Complex Reasoning (Qwen 3.6 35B)     |   |
|   |  [@Codebase Index]   --> Vector Embeddings (nomic-embed-text) |   |
|   +-------------------------------+-------------------------------+   |
+-----------------------------------|-----------------------------------+
                                    |
            HTTP / HTTPS via LAN or Tailscale Mesh (Port 11434)
                                    |
+-----------------------------------v-----------------------------------+
|  Linux Server / Homelab Node (Ubuntu 24.04 LTS / Debian 12)           |
|                                                                       |
|   +---------------------------------------------------------------+   |
|   | Docker Container: Ollama Inference Engine                     |   |
|   |                                                               |   |
|   |   - NVIDIA Container Toolkit (CUDA Passthrough)               |   |
|   |   - OLLAMA_KEEP_ALIVE=24h (Instant Model Response)            |   |
|   |   - OLLAMA_ORIGINS="*" (Allowed Web/IDE Extensions)           |   |
|   |   - Persistent Storage: /root/.ollama (Models Cache)          |   |
|   +-------------------------------+-------------------------------+   |
|                                   |                                   |
|   +-------------------------------v-------------------------------+   |
|   | Host Hardware: GPU VRAM (RTX 4090 / A5000 / RTX 3090)         |   |
|   +---------------------------------------------------------------+   |
+-----------------------------------------------------------------------+

Prerequisites & VRAM Allocation

Before launching the stack, verify that your server satisfies the minimum hardware and software prerequisites:

  • Linux OS: Ubuntu 22.04/24.04 LTS, Debian 12, or Rocky Linux 9 with root/sudo privileges.
  • NVIDIA Drivers: Version 535.xx or newer installed on the host. Verify with nvidia-smi.
  • NVIDIA Container Toolkit: Configured as the default runtime in /etc/docker/daemon.json.
  • VRAM Capacity:
    • 8 GB VRAM: Suitable for qwen3.6:7b-coder (quantized 4-bit) for both autocomplete and chat.
    • 16 GB VRAM: Capable of hosting qwen3.6:14b-coder alongside an embedding model.
    • 24 GB+ VRAM: The sweet spot for running qwen3.6:35b-coder for deep architectural refactoring and qwen3.6:7b for sub-100ms autocomplete simultaneously.

Check your NVIDIA container runtime readiness on the host before proceeding:

# Verify host driver and CUDA capabilities
nvidia-smi

# Test Docker container GPU passthrough
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

Step 1: Production Docker Compose Configuration

Create a dedicated directory on your server to house the Ollama deployment:

mkdir -p /opt/ollama-continue && cd /opt/ollama-continue

Create the docker-compose.yml file. We configure specific environment parameters crucial for coding assistants: OLLAMA_ORIGINS must allow requests originating from IDE browser runtimes (VS Code web views), OLLAMA_KEEP_ALIVE prevents the model from unloading between pauses in typing, and OLLAMA_NUM_PARALLEL enables concurrent requests so autocomplete queries do not block conversational chat.

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama-service
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ollama_models:/root/.ollama
    environment:
      - OLLAMA_KEEP_ALIVE=24h
      - OLLAMA_ORIGINS=vscode-webview://*,vscode-file://*,chrome-extension://*,http://localhost:*,http://127.0.0.1:*
      - OLLAMA_NUM_PARALLEL=4
      - OLLAMA_MAX_LOADED_MODELS=2
      - OLLAMA_FLASH_ATTENTION=1
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

volumes:
  ollama_models:
    name: ollama_models_storage

Launch the container in detached mode:

docker compose up -d

Verify that Ollama initialized and identified your GPU properly:

docker logs -f ollama-service

Step 2: Pulling State-of-the-Art Qwen 3.6 & Embedding Models

Modern coding workflows require three distinct model roles:

  1. Chat & Instruction Model: High parameter count for architectural analysis, writing complex algorithms, and reviewing pull requests (e.g., qwen3.6:35b-coder or qwen3.6:14b-coder).
  2. Tab Autocomplete (Fill-in-the-Middle – FIM): Ultra-fast, lower-latency model optimized specifically for completing the current line or function prefix/suffix (e.g., qwen3.6:7b-coder).
  3. Embeddings Model: Lightweight transformer model used by Continue to vectorize and search your entire workspace repository for contextual retrieval (e.g., nomic-embed-text).

Execute the following commands to pull the models into your persistent Docker volume:

# Pull the primary chat model (Qwen 3.6 Coder)
docker exec -it ollama-service ollama pull qwen3.6:35b-coder

# Pull the high-speed FIM autocomplete model
docker exec -it ollama-service ollama pull qwen3.6:7b-coder

# Pull the local embedding model for codebase vector search
docker exec -it ollama-service ollama pull nomic-embed-text

Inspect the pulled models to ensure proper quantization and storage:

docker exec -it ollama-service ollama list

Step 3: Configuring the Continue.dev Extension in VS Code

Install the Continue extension from the Visual Studio Code Marketplace (or Open VSX Registry in VSCodium). Once installed, click the Continue icon in the activity bar, open settings (gear icon), or edit the configuration file directly located at:

  • Linux / macOS: ~/.continue/config.yaml
  • Windows: %USERPROFILE%\.continue\config.yaml

Paste the following production configuration into your config.yaml. If your Ollama server is hosted on a separate machine across your local network or VPN, replace http://localhost:11434 with your server’s LAN IP or Tailscale domain name (e.g., http://192.168.1.150:11434):

name: Local-SelfHosted-AI
models:
  - name: "Qwen 3.6 Coder 35B"
    provider: "ollama"
    model: "qwen3.6:35b-coder"
    apiBase: "http://192.168.1.150:11434"
    roles:
      - "chat"
      - "edit"
      - "apply"
    requestOptions:
      timeout: 120

tabAutocompleteModel:
  name: "Qwen 3.6 Coder 7B Autocomplete"
  provider: "ollama"
  model: "qwen3.6:7b-coder"
  apiBase: "http://192.168.1.150:11434"

embeddingsProvider:
  provider: "ollama"
  model: "nomic-embed-text"
  apiBase: "http://192.168.1.150:11434"

tabAutocompleteOptions:
  useCache: true
  debounceDelay: 250
  maxPromptTokens: 1024

customCommands:
  - name: test
    prompt: "Write a comprehensive test suite covering edge cases for the selected code using modern testing frameworks."
    description: "Generate unit tests"
  - name: refactor
    prompt: "Refactor the selected code for readability, performance, and adherence to clean code principles without altering behavior."
    description: "Refactor code"

Save the file. Continue will automatically hot-reload the configuration. Test the setup by pressing Ctrl + L (or Cmd + L on macOS) to open the chat window, type Explain the current function, and verify that inference streaming begins immediately from your local container.

Step 4: Indexing Codebases with Local RAG (@codebase)

One of the most potent capabilities of Continue.dev is the @codebase context provider. By leveraging your locally hosted nomic-embed-text model, Continue scans your project workspace, generates embeddings for code chunks, and stores them in a local SQLite-backed LanceDB vector store on your machine.

To initialize repository indexing:

  1. Open a repository in your IDE.
  2. In the Continue chat prompt, type @codebase How does our authentication middleware handle JWT validation?.
  3. Continue will perform semantic search across your entire codebase, inject relevant file snippets into the prompt, and pass the context to Qwen 3.6 for an accurate, repository-aware answer.

To exclude temporary artifacts, build directories, and environment secrets from being indexed, ensure your project contains an .ignore or .gitignore file containing entries such as node_modules/, dist/, venv/, and .env.

Step 5: Security Hardening & Network Isolation

By default, Ollama binds to 0.0.0.0:11434 without built-in authentication. In a multi-user corporate or homelab environment, exposing this port directly to an untrusted subnet invites unauthorized resource exhaustion. Implement the following hardening controls:

  • Tailscale Private Mesh: Bind the host port strictly to the Tailscale interface (e.g., 100.x.y.z:11434:11434 in your Compose file) so only authenticated devices can communicate with Ollama.
  • Caddy Reverse Proxy with Basic Auth / API Keys: Front the container with Caddy or Nginx to enforce SSL encryption and Bearer token headers if you require remote internet access.
  • Firewall Ingress Rules: Use UFW or iptables to restrict traffic on port 11434 to explicit workstation IP addresses:
# Deny public ingress on port 11434
sudo ufw deny 11434/tcp

# Allow specific developer machine
sudo ufw allow from 192.168.1.50 to any port 11434 proto tcp

Troubleshooting Common Production Issues

1. GPU Not Detected / Falling Back to Slow CPU Inference

Symptoms: Token generation is painfully slow (1–3 tokens/second), and nvidia-smi reveals 0% GPU utilization during active prompts.

Resolution: Ensure the host has the NVIDIA Container Toolkit configured. Edit /etc/docker/daemon.json to register the runtime:

{
  "default-runtime": "nvidia",
  "runtimes": {
    "nvidia": {
      "path": "nvidia-container-runtime",
      "runtimeArgs": []
    }
  }
}

Restart the Docker daemon with sudo systemctl restart docker and recreate your containers with docker compose down && docker compose up -d.

2. Continue.dev Shows “Failed to fetch” or CORS Rejection

Symptoms: The Continue panel displays a red warning reading Connection failed to http://server:11434 or browser developer tools log CORS header 'Access-Control-Allow-Origin' missing.

Resolution: The Ollama daemon rejects requests from unknown webview origins unless explicitly authorized. Ensure your docker-compose.yml contains OLLAMA_ORIGINS="*" or includes vscode-webview://*. Apply the change and restart the container:

docker compose down && docker compose up -d

3. Tab-Autocomplete Delays or Out-of-Memory (OOM) Container Crash

Symptoms: When typing in your editor, tab completions lag by 2–4 seconds, or the Ollama container abruptly exits with status code 137 (OOM killed).

Resolution: Concurrently loading two large models (a 35B chat model and a 7B autocomplete model) requires at least 24 GB VRAM. If your GPU has 12–16 GB VRAM, configure OLLAMA_MAX_LOADED_MODELS=1 or downgrade your chat model to qwen3.6:14b-coder. Additionally, adjust the debounceDelay in config.yaml to 350ms to prevent rapid keystrokes from queueing unnecessary completion requests.

Conclusion & Key Takeaways

Deploying Continue.dev with Ollama and Qwen 3.6 inside Docker Compose provides the ultimate balance between cutting-edge AI coding acceleration and uncompromising data privacy. By hosting your own open-weights models, your codebase remains strictly within your physical or virtual perimeter while your engineering workflow gains instant tab autocompletions, codebase-wide semantic search, and autonomous refactoring.

To take your private AI infrastructure even further, consider integrating a LiteLLM proxy in Docker for centralized team key management, or connect Model Context Protocol (MCP) servers to grant your assistant deep access to local databases, git branches, and continuous integration logs.