Deploying Longhorn on K3s for Distributed Persistent Block Storage

Deploy Longhorn on K3s for distributed, highly available persistent block storage. Step-by-step guide with iSCSI host preparation, Helm deployment, Traefik UI ingress, and automated S3 backups.

When running stateful workloads—such as relational databases, private document stores, or persistent container volumes—inside a K3s Kubernetes cluster, data persistence quickly becomes the primary operational bottleneck. By default, K3s includes Rancher’s lightweight local-path provisioner. While local-path is blazingly fast for single-node development, it binds storage volumes strictly to the local directory of the specific host where the pod initially spawned. If that physical node reboots, suffers hardware degradation, or undergoes scheduled maintenance, your pod cannot be rescheduled on another node without losing access to its data.

To achieve true production-grade resilience, modern cloud-native architectures require distributed, replicated block storage. Enter Longhorn: a CNCF-incubated, lightweight, open-source distributed block storage system designed specifically for Kubernetes. Longhorn transforms local node disks into enterprise-class storage pools, creating synchronous block-level replicas across multiple nodes while offering incremental snapshots, automated offsite backups to S3-compatible object storage, and an intuitive management dashboard.

In this comprehensive hands-on guide, you will learn how to prepare your K3s nodes with required kernel modules and dependencies, install and configure Longhorn, securely expose the Longhorn UI behind Traefik Ingress with TLS, provision highly available PersistentVolumeClaims (PVCs), and automate recurring backups to S3.

Why Choose Longhorn for K3s Storage?

Unlike monolithic SAN/NAS storage appliances or heavy enterprise distributed filesystems like Ceph (Rook), Longhorn is tailored specifically for resource-conscious Kubernetes environments such as edge gateways, homelabs, and private clouds. According to the Official Longhorn Documentation and the K3s Storage Guide, Longhorn delivers several distinct operational advantages:

  • Microservice-Centric Block Storage: Each volume runs its own dedicated storage controller and replica engines inside lightweight containers, eliminating cluster-wide single points of failure.
  • Synchronous Multi-Node Replication: Volumes can be configured with 2 or 3 synchronous replicas across separate physical nodes. If a worker node crashes, Kubernetes seamlessly re-attaches the volume on another healthy node in seconds.
  • Incremental S3 & NFS Backups: Longhorn takes thin-provisioned snapshots and uploads incremental block deltas directly to AWS S3, MinIO, or Cloudflare R2, similar to the deduplicated strategy we explored in our Docker Restic backup tutorial.
  • Zero-Downtime Volume Expansion: Resize PersistentVolumeClaims dynamically without unmounting active workloads or stopping database processes.
  • Cross-Cluster Disaster Recovery: Standby volumes can continuously poll S3 backup targets to enable rapid cross-cluster recovery.

Prerequisites & Host Preparation

Longhorn operates at the Linux block device level using standard iSCSI and container storage interface (CSI) protocols. Before installing Longhorn, every participating node in your K3s cluster must have the required kernel modules and storage utilities installed.

Execute the following commands on all K3s server and agent nodes:

# Update package repositories and install required storage utilities
sudo apt update && sudo apt install -y open-iscsi nfs-common util-linux cryptsetup jq curl

# Ensure the iSCSI service is enabled and actively running
sudo systemctl enable --now iscsid

# Verify that the iscsi_tcp kernel module is loaded
sudo modprobe iscsi_tcp
echo "iscsi_tcp" | sudo tee -a /etc/modules-load.d/iscsi.conf

To confirm that your host nodes meet every single architectural requirement (including mount propagation, Linux kernel version, and cgroups configuration), run the official Longhorn environment validation script using kubectl:

curl -sSfL https://raw.githubusercontent.com/longhorn/longhorn/master/scripts/environment_check.sh | bash

Every check must report [PASS] across your nodes before proceeding.

Step 1: Installing Longhorn via Helm

While Longhorn can be deployed via raw Kubernetes manifests, deploying through Helm provides predictable versioning, simple upgrades, and fine-grained parameter overrides.

First, add the official Longhorn Helm repository on your management workstation or control-plane node:

helm repo add longhorn https://charts.longhorn.io
helm repo update

Next, create a custom values file named longhorn-values.yaml to tune replication parameters for your cluster. For two-node or three-node clusters, setting default replicas to 2 or 3 ensures balanced redundancy:

cat << 'EOF' > longhorn-values.yaml
defaultSettings:
  defaultDataPath: "/var/lib/longhorn"
  defaultReplicaCount: 2
  storageOverProvisioningPercentage: 200
  storageMinimalAvailablePercentage: 15
  guaranteedInstanceManagerCPU: 10
  replicaAutoBalance: "least-effort"

persistence:
  defaultClass: true
  defaultClassReplicaCount: 2
  reclaimPolicy: "Delete"

csi:
  kubeletRootDir: "/var/lib/kubelet"
EOF

Install Longhorn into its dedicated longhorn-system namespace:

helm install longhorn longhorn/longhorn \
  --namespace longhorn-system \
  --create-namespace \
  -f longhorn-values.yaml

Monitor the rollout until all Longhorn CSI plugins, engine managers, and controllers reach Running status:

kubectl get pods -n longhorn-system -w

Once deployed, verify that Longhorn registered its default StorageClass:

kubectl get storageclass

You should see longhorn (default) alongside K3s’s built-in local-path provisioner.

Step 2: Exposing and Securing the Longhorn Dashboard

Longhorn provides an intuitive web interface for inspecting storage disks, real-time replica synchronization, volume throughput, and snapshot management. Because the dashboard has no built-in user authentication out of the box, exposing it directly to the internet is a severe vulnerability. We will secure it using Traefik Basic Authentication and Let’s Encrypt SSL.

First, generate an encrypted password hash using htpasswd:

# Install apache2-utils if htpasswd is not present
sudo apt install -y apache2-utils

# Create an auth secret with username 'admin'
htpasswd -nb admin "YourSecureClusterPasswordHere" > auth

Create the Kubernetes secret and Traefik Middleware inside the longhorn-system namespace:

cat << 'EOF' | kubectl apply -f -
apiVersion: v1
kind: Secret
metadata:
  name: basic-auth-secret
  namespace: longhorn-system
type: Opaque
data:
  auth: $(cat auth | base64 -w 0)
---
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
  name: longhorn-auth
  namespace: longhorn-system
spec:
  basicAuth:
    secret: basic-auth-secret
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: longhorn-ingress
  namespace: longhorn-system
  annotations:
    traefik.ingress.kubernetes.io/router.middlewares: longhorn-system-longhorn-auth@kubernetescrd
    cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
  ingressClassName: traefik
  tls:
    - hosts:
        - longhorn.yourdomain.com
      secretName: longhorn-dashboard-tls
  rules:
    - host: longhorn.yourdomain.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: longhorn-frontend
                port:
                  number: 80
EOF

Navigating to https://longhorn.yourdomain.com now prompts for HTTP Basic Authentication before granting access to your cluster storage topology.

Step 3: Deploying a Stateful Workload with Replicated Storage

Let us test dynamic volume provisioning by deploying a production-ready PostgreSQL database with a 2-replica Persistent Volume Claim backed by Longhorn.

Create a dedicated namespace and deployment manifest:

cat << 'EOF' | kubectl apply -f -
apiVersion: v1
kind: Namespace
metadata:
  name: storage-demo
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: postgres-data-pvc
  namespace: storage-demo
spec:
  accessModes:
    - ReadWriteOnce
  storageClassName: longhorn
  resources:
    requests:
      storage: 20Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: postgres-demo
  namespace: storage-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: postgres-demo
  template:
    metadata:
      labels:
        app: postgres-demo
    spec:
      containers:
        - name: postgres
          image: postgres:16-alpine
          env:
            - name: POSTGRES_DB
              value: "appdb"
            - name: POSTGRES_USER
              value: "dbadmin"
            - name: POSTGRES_PASSWORD
              value: "SuperSecretPassw0rd!"
            - name: PGDATA
              value: "/var/lib/postgresql/data/pgdata"
          ports:
            - containerPort: 5432
              name: postgres
          volumeMounts:
            - name: db-storage
              mountPath: /var/lib/postgresql/data
      volumes:
        - name: db-storage
          persistentVolumeClaim:
            claimName: postgres-data-pvc
EOF

Verify that Longhorn dynamically provisioned and bound the Persistent Volume:

kubectl get pvc,pv -n storage-demo

Inspect the Longhorn web dashboard. You will see postgres-data-pvc listed as Healthy, with two synchronous block replicas distributed across separate physical worker nodes.

Step 4: Automating S3 Snapshots and Disaster Recovery

Having synchronized replicas protects against single-node drive failures, but it does not protect against catastrophic datacenter outages, accidental table drops, or ransomware. Longhorn includes native S3 backup integration that takes incremental, deduplicated snapshots.

Create an AWS or MinIO S3 credentials secret in the longhorn-system namespace:

cat << 'EOF' | kubectl apply -f -
apiVersion: v1
kind: Secret
metadata:
  name: s3-backup-secret
  namespace: longhorn-system
type: Opaque
stringData:
  AWS_ACCESS_KEY_ID: "YOUR_S3_ACCESS_KEY"
  AWS_SECRET_ACCESS_KEY: "YOUR_S3_SECRET_KEY"
  AWS_ENDPOINTS: "https://s3.eu-central-1.amazonaws.com"
EOF

In the Longhorn UI, navigate to Settings > General > Backup Target and configure:

  • Backup Target: s3://your-cluster-backups@eu-central-1/k3s-longhorn
  • Backup Target Credential Secret: s3-backup-secret

Now, define an automated RecurringJob that takes a local snapshot every 6 hours and pushes an offsite backup to S3 every night at 02:00 UTC:

cat << 'EOF' | kubectl apply -f -
apiVersion: longhorn.io/v1beta2
kind: RecurringJob
metadata:
  name: daily-s3-backup
  namespace: longhorn-system
spec:
  cron: "0 2 * * *"
  task: "backup"
  retain: 14
  concurrency: 2
  labels:
    backup-policy: "daily"
EOF

Step 5: Testing Node Eviction and Failover Resilience

To verify that our distributed storage architecture truly prevents downtime when hardware fails, let us simulate a sudden node outage by cordoning and draining the active worker node hosting the PostgreSQL pod:

# Identify the node currently running the database
kubectl get pods -n storage-demo -o wide

# Cordon and drain the node
kubectl drain <worker-node-name> --ignore-daemonsets --delete-emptydir-data --force

Observe the Kubernetes scheduler and Longhorn in real time:

  1. The PostgreSQL pod on the drained node enters Terminating.
  2. Kubernetes reschedules the pod to the surviving worker node.
  3. The Longhorn CSI driver detaches the volume from the old node and attaches it to the surviving node in approximately 8 to 15 seconds.
  4. The PostgreSQL container mounts /var/lib/postgresql/data from the surviving local replica and resumes serving database queries without data loss.

Once your maintenance simulation is complete, uncordon the node:

kubectl uncordon <worker-node-name>

Longhorn automatically detects that the node is online again, recalculates replica health, and synchronizes any missing block deltas in the background.

Troubleshooting Common Longhorn & K3s Issues

Symptom Root Cause Resolution
Volume stuck in Attaching Missing or inactive iscsid daemon on worker node. Install open-iscsi on the host and run sudo systemctl restart iscsid.
Replica state Degraded Cluster lacks enough distinct nodes or disks to satisfy replica count. Adjust numberOfReplicas to match your active worker node count or add worker nodes.
Disk Out of Space /var/lib/longhorn shared with OS root disk filled up. Attach a secondary dedicated NVMe/SSD disk and configure it via Longhorn Node Disk settings.

Summary & Next Steps

Deploying Longhorn on K3s elevates your Kubernetes infrastructure from an experimental single-node sandbox into a robust, fault-tolerant cluster. With synchronous block-level replication, intuitive web management, and automated offsite S3 backups, your databases and mission-critical containers can survive physical node outages with zero administrative panic.

To maximize Longhorn’s multi-node replication capabilities, your next architectural milestone is expanding your cluster topology by joining dedicated worker agent nodes to your control plane.