How to Deploy Dragonfly with Docker Compose: High-Throughput Redis & Memcached Drop-In Replacement

Dragonfly In-Memory Cache mit Docker Compose bereitstellen - Cloud-Infrastruktur-Ingenieurin im Serverraum vor Datenbank-Monitoring-Dashboards
How to Deploy Dragonfly with Docker Compose: High-Throughput Redis & Memcached Drop-In Replacement 3

In modern web infrastructure, microservices, and distributed applications, in-memory caching layers dictate whether your systems scale smoothly under burst traffic or collapse under database bottlenecks. For more than a decade, Redis has served as the undisputed industry standard for key-value storage, session persistence, and pub/sub message brokers. However, the architectural foundation of Redis dates back to 2009—an era when single-core processing was the baseline and multi-core server CPUs were in their infancy.

As modern cloud virtual machines and bare-metal servers scale to 32, 64, or 128 hardware threads, traditional Redis struggles to exploit that parallelism. Because Redis operates primarily on a single-threaded event loop, scaling horizontally requires complex Redis Cluster topologies, proxy layers like Envoy or Twemproxy, or running multiple siloed Redis instances per host. Furthermore, Redis’s traditional persistence mechanism relies on Linux’s fork() system call, which can trigger severe memory amplification and tail-latency spikes under heavy write workloads.

Enter Dragonfly: an open-source, in-memory data store engineered from scratch in C++ to act as a 100% wire-compatible drop-in replacement for both Redis and Memcached. Dragonfly leverages a shared-nothing, multi-threaded architecture with asynchronous I/O (via Linux io_uring) and an innovative lock-free hash table called Dashtable. The result is up to 25x higher throughput and sub-millisecond P99 tail latency while consuming up to 30% less RAM than traditional Redis on identical hardware. In this guide, you will learn how to deploy a hardened, production-grade Dragonfly instance using Docker Compose, complete with persistence snapshotting, authentication, kernel optimizations, and performance benchmarking.

Architecture Comparison: Redis vs. Dragonfly

Understanding why Dragonfly outperforms Redis requires examining how both engines handle memory, concurrency, and persistent snapshots on modern Linux kernels.

+-----------------------------------------------------------------------------------+
|                            TRADITIONAL REDIS ENGINE                               |
|                                                                                   |
|  [ Client 1 ]   [ Client 2 ]   [ Client 3 ]   [ Client 4 ]                        |
|        \              |              |              /                             |
|         v             v              v             v                              |
|  +-----------------------------------------------------------------------------+  |
|  | Single-Threaded Event Loop (Single CPU Core Saturated)                      |  |
|  | - Other 15–63 CPU cores idle on modern hardware                             |  |
|  | - Snapshotting (BGSAVE) uses fork() -> Copy-on-Write memory doubling risk   |  |
|  +-----------------------------------------------------------------------------+  |
+-----------------------------------------------------------------------------------+

+-----------------------------------------------------------------------------------+
|                         DRAGONFLY SHARED-NOTHING ENGINE                           |
|                                                                                   |
|  [ Client 1 ]   [ Client 2 ]   [ Client 3 ]   [ Client 4 ]                        |
|        |              |              |              |                             |
|        v              v              v              v                             |
|  +------------+ +------------+ +------------+ +------------+                      |
|  | Thread 1   | | Thread 2   | | Thread 3   | | Thread 4   | (N Cores Scaled)     |
|  | io_uring   | | io_uring   | | io_uring   | | io_uring   |                      |
|  +------------+ +------------+ +------------+ +------------+                      |
|        |              |              |              |                             |
|        v              v              v              v                             |
|  +-----------------------------------------------------------------------------+  |
|  | Dashtable: Partitioned Lock-Free In-Memory Hash Table                       |  |
|  | - Zero global mutex contention across threads                               |  |
|  | - Zero-copy checkpointing without fork() latency spikes                    |  |
|  | - Transparent compression & up to 30% lower memory overhead                |  |
|  +-----------------------------------------------------------------------------+  |
+-----------------------------------------------------------------------------------+

Dragonfly divides the global keyspace across CPU threads into isolated shards. Each thread executes inside an autonomous event loop powered by io_uring, preventing costly context switching and cross-thread lock contention. When persisting snapshots to disk, Dragonfly uses an asynchronous checkpointing algorithm that captures memory states inline without invoking Linux fork(), eliminating the notorious out-of-memory crashes caused by copy-on-write page duplication during heavy writes.

Prerequisites & Linux Kernel Tuning

To maximize Dragonfly’s throughput and prevent memory allocation stalls on your Docker host (Ubuntu 22.04/24.04 or Debian 12), apply these essential kernel sysctl parameters before launching the container.

1. Host Memory Overcommit and File Descriptors

Set Linux memory overcommit to heuristic allocation mode (1) and ensure socket connection backlogs can absorb massive concurrent connection spikes:

# Apply temporary runtime settings
sudo sysctl vm.overcommit_memory=1
sudo sysctl -w fs.file-max=2097152
sudo sysctl -w net.core.somaxconn=65535

# Persist settings in /etc/sysctl.d/99-dragonfly.conf
sudo tee /etc/sysctl.d/99-dragonfly.conf <<EOF
vm.overcommit_memory = 1
fs.file-max = 2097152
net.core.somaxconn = 65535
EOF

sudo sysctl --system

2. Project Directory Structure

Create a dedicated workspace directory with persistent storage mounts for data dumps and snapshot backups:

mkdir -p ~/dragonfly-stack/data
cd ~/dragonfly-stack

Step 1: Environment Configuration

Create an environment file ~/dragonfly-stack/.env to store credentials and runtime memory constraints:

# Dragonfly Configuration
DRAGONFLY_PORT=6379
DRAGONFLY_PASSWORD=DragonflySecureCacheMasterKey2026!
DRAGONFLY_MAX_MEMORY=4GB
DRAGONFLY_THREADS=4
DRAGONFLY_SNAPSHOT_INTERVAL=3600

Step 2: Production Docker Compose Specification

Create ~/dragonfly-stack/docker-compose.yml. The configuration below passes explicit tuning flags to the Dragonfly binary, mounts persistent storage for RDB/DFS snapshots, configures container resource ceilings, and implements a native Redis-protocol healthcheck:

services:
  dragonfly:
    image: docker.dragonflydb.io/dragonflydb/dragonfly:latest
    container_name: dragonfly
    restart: unless-stopped
    ulimits:
      memlock: -1
      nofile:
        soft: 65535
        hard: 65535
    ports:
      - "${DRAGONFLY_PORT}:6379"
    environment:
      - TZ=UTC
    command:
      - dragonfly
      - --logtostderr
      - --requirepass=${DRAGONFLY_PASSWORD}
      - --maxmemory=${DRAGONFLY_MAX_MEMORY}
      - --threads=${DRAGONFLY_THREADS}
      - --dbfilename=dump.rdb
      - --dir=/data
      - --save_schedule=*:${DRAGONFLY_SNAPSHOT_INTERVAL}
      - --admin_port=9999
      - --proclist=true
    volumes:
      - ./data:/data
    networks:
      - cache_net
    healthcheck:
      test: ["CMD-SHELL", "redis-cli -a ${DRAGONFLY_PASSWORD} ping | grep -q PONG"]
      interval: 10s
      timeout: 3s
      retries: 3
      start_period: 5s
    deploy:
      resources:
        limits:
          cpus: '${DRAGONFLY_THREADS}'
          memory: 4.5G
        reservations:
          cpus: '1.0'
          memory: 1G

networks:
  cache_net:
    driver: bridge

Key Parameter Breakdown:

  • --requirepass: Enforces authentication across all client commands. Wire-compatible with AUTH <password>.
  • --maxmemory: Caps active RAM consumption. When reached, Dragonfly applies an intelligent eviction policy without allocating unbounded memory.
  • --threads: Specifies the number of execution threads. On multi-core servers, set this to the number of physical CPU cores allocated to caching workloads.
  • --dir and --dbfilename: Defines the location and filename for persistent snapshot dumps on host disk.
  • --admin_port=9999: Exposes an internal HTTP administrative port providing Prometheus-compatible metrics (/metrics) and runtime diagnostics.
  • ulimits: nofile: 65535: Raises the maximum open file descriptors so Dragonfly can maintain tens of thousands of concurrent client connections without dropping sockets.

Step 3: Starting and Validating the Container

Start Dragonfly in detached mode using Docker Compose:

docker compose up -d

Check the container status and confirm that the health check status is healthy:

docker compose ps

Review the startup logs to verify CPU hardware detection and thread pool initialization:

docker compose logs dragonfly

You should see Dragonfly report its thread count, memory limit, and listener bindings:

I2026-10-08 00:15:22 server_family.cc:510] Starting dragonfly on :6379 with 4 threads
I2026-10-08 00:15:22 server_family.cc:512] Max memory: 4.00 GiB
I2026-10-08 00:15:22 server_family.cc:538] Running in Redis 7.2 compatibility mode

Step 4: Functional Testing with redis-cli

Because Dragonfly implements the standard Redis Serialization Protocol (RESP2 and RESP3), any standard Redis client library (such as Python redis-py, Node.js ioredis, Go go-redis, or PHP Predis) connects immediately without modifications.

Connect to Dragonfly using the embedded redis-cli tool inside the container:

docker exec -it dragonfly redis-cli -a DragonflySecureCacheMasterKey2026!

Execute standard key-value, pipeline, and telemetry commands:

127.0.0.1:6379> PING
PONG

127.0.0.1:6379> SET user:1001:session "eyJhbGciOiJIUzI1NiJ9..." EX 3600
OK

127.0.0.1:6379> GET user:1001:session
"eyJhbGciOiJIUzI1NiJ9..."

127.0.0.1:6379> MSET k1 "v1" k2 "v2" k3 "v3"
OK

127.0.0.1:6379> KEYS *
1) "user:1001:session"
2) "k1"
3) "k2"
4) "k3"

127.0.0.1:6379> INFO server
# Server
dragonfly_version:1.26.0
redis_version:7.2.4
os:Linux 6.8.0-45-generic x86_64
arch_bits:64
multiplexing_api:io_uring
process_id:1
tcp_port:6379
uptime_in_seconds:120

Step 5: Benchmarking Throughput and Latency

To measure the raw throughput benefits of Dragonfly’s multi-threaded Dashtable architecture, execute a standardized benchmark directly against the container using redis-benchmark:

docker exec -it dragonfly redis-benchmark \
  -a DragonflySecureCacheMasterKey2026! \
  -h 127.0.0.1 \
  -p 6379 \
  -c 50 \
  -n 500000 \
  -d 128 \
  -t set,get \
  -q

On a standard 4-core virtual server, Dragonfly easily achieves over 180,000 to 250,000 requests per second with P99 latency remaining consistently under 0.4 milliseconds:

SET: 215424.38 requests per second, p50=0.175 msec, p99=0.383 msec
GET: 246913.58 requests per second, p50=0.151 msec, p99=0.311 msec

By comparison, a single-threaded Redis instance running on identical hardware bottlenecks around 80,000 to 100,000 QPS as CPU core 0 pegs at 100% utilization.

Step 6: Triggering and Verifying Persistent Snapshots

Dragonfly supports on-demand background snapshots via the traditional BGSAVE command:

docker exec -it dragonfly redis-cli -a DragonflySecureCacheMasterKey2026! BGSAVE

Verify that the dump file has been written to the host directory (~/dragonfly-stack/data):

ls -lh ~/dragonfly-stack/data/dump.rdb

Because Dragonfly writes standard RDB-format dumps, you can take a snapshot generated by Dragonfly and restore it directly into a standard Redis instance—or restore an existing Redis RDB dump directly into Dragonfly upon container boot.

Troubleshooting Common Issues

1. “OOM command not allowed when used memory > ‘maxmemory'”

Cause: Dragonfly has ingested data up to the --maxmemory threshold and cannot evict keys because existing keys have no expiration TTL or the dataset consists entirely of non-volatile keys.

Solution: Check memory usage via redis-cli -a <PASS> INFO memory. Either increase DRAGONFLY_MAX_MEMORY in your .env file (e.g., from 4GB to 8GB) or configure an aggressive eviction policy by adding --maxmemory_policy=allkeys-lru to the command block in docker-compose.yml.

2. “io_uring initialization failed: Operation not permitted”

Cause: Older Docker engines, hardened AppArmor profiles, or restricted Docker seccomp profiles may block the io_uring_setup system call. This is common on legacy host kernels (Linux 5.4 or older).

Solution: If upgrading your Linux host kernel is not immediately possible, force Dragonfly to fall back to the standard epoll I/O multiplexer by adding --force_epoll=true to the command: section in docker-compose.yml. On modern kernels (Linux 5.15+), ensure Docker’s seccomp filter allows io_uring by updating container capabilities.

3. Client Connection Errors: “Could not connect to Redis: Connection refused” or Socket Drops

Cause: High-concurrency load testing exhaustively consumes ephemeral ports or exceeds the container’s file descriptor limit (nofile).

Solution: Verify that ulimits.nofile.hard and ulimits.nofile.soft are both explicitly set to 65535 in docker-compose.yml. On the host operating system, verify that sysctl net.core.somaxconn is set to 65535 to prevent Linux network stack queue drops during micro-bursts.

Conclusion

Dragonfly represents the natural evolution of in-memory key-value caching for the multi-core cloud era. By replacing single-threaded event loops with an asynchronous, shared-nothing C++ engine, Dragonfly extracts maximum hardware efficiency from modern server CPUs while remaining fully compatible with existing Redis client libraries and tools. Deploying Dragonfly with Docker Compose gives you instant access to massive throughput gains, sub-millisecond tail latencies, robust RDB snapshotting, and lower memory overhead—without rewriting a single line of application code.