When running stateful workloads—such as relational databases, private document stores, or persistent container volumes—inside a K3s Kubernetes cluster, data persistence quickly becomes the primary operational bottleneck. By default, K3s includes Rancher’s lightweight local-path provisioner. While local-path is blazingly fast for single-node development, it binds storage volumes strictly to the local directory of the specific host where the pod initially spawned. If that physical node reboots, suffers hardware degradation, or undergoes scheduled maintenance, your pod cannot be rescheduled on another node without losing access to its data.
To achieve true production-grade resilience, modern cloud-native architectures require distributed, replicated block storage. Enter Longhorn: a CNCF-incubated, lightweight, open-source distributed block storage system designed specifically for Kubernetes. Longhorn transforms local node disks into enterprise-class storage pools, creating synchronous block-level replicas across multiple nodes while offering incremental snapshots, automated offsite backups to S3-compatible object storage, and an intuitive management dashboard.
In this comprehensive hands-on guide, you will learn how to prepare your K3s nodes with required kernel modules and dependencies, install and configure Longhorn, securely expose the Longhorn UI behind Traefik Ingress with TLS, provision highly available PersistentVolumeClaims (PVCs), and automate recurring backups to S3.
Why Choose Longhorn for K3s Storage?
Unlike monolithic SAN/NAS storage appliances or heavy enterprise distributed filesystems like Ceph (Rook), Longhorn is tailored specifically for resource-conscious Kubernetes environments such as edge gateways, homelabs, and private clouds. According to the Official Longhorn Documentation and the K3s Storage Guide, Longhorn delivers several distinct operational advantages:
- Microservice-Centric Block Storage: Each volume runs its own dedicated storage controller and replica engines inside lightweight containers, eliminating cluster-wide single points of failure.
- Synchronous Multi-Node Replication: Volumes can be configured with 2 or 3 synchronous replicas across separate physical nodes. If a worker node crashes, Kubernetes seamlessly re-attaches the volume on another healthy node in seconds.
- Incremental S3 & NFS Backups: Longhorn takes thin-provisioned snapshots and uploads incremental block deltas directly to AWS S3, MinIO, or Cloudflare R2, similar to the deduplicated strategy we explored in our Docker Restic backup tutorial.
- Zero-Downtime Volume Expansion: Resize PersistentVolumeClaims dynamically without unmounting active workloads or stopping database processes.
- Cross-Cluster Disaster Recovery: Standby volumes can continuously poll S3 backup targets to enable rapid cross-cluster recovery.
Prerequisites & Host Preparation
Longhorn operates at the Linux block device level using standard iSCSI and container storage interface (CSI) protocols. Before installing Longhorn, every participating node in your K3s cluster must have the required kernel modules and storage utilities installed.
Execute the following commands on all K3s server and agent nodes:
# Update package repositories and install required storage utilities
sudo apt update && sudo apt install -y open-iscsi nfs-common util-linux cryptsetup jq curl
# Ensure the iSCSI service is enabled and actively running
sudo systemctl enable --now iscsid
# Verify that the iscsi_tcp kernel module is loaded
sudo modprobe iscsi_tcp
echo "iscsi_tcp" | sudo tee -a /etc/modules-load.d/iscsi.conf
To confirm that your host nodes meet every single architectural requirement (including mount propagation, Linux kernel version, and cgroups configuration), run the official Longhorn environment validation script using kubectl:
curl -sSfL https://raw.githubusercontent.com/longhorn/longhorn/master/scripts/environment_check.sh | bash
Every check must report [PASS] across your nodes before proceeding.
Step 1: Installing Longhorn via Helm
While Longhorn can be deployed via raw Kubernetes manifests, deploying through Helm provides predictable versioning, simple upgrades, and fine-grained parameter overrides.
First, add the official Longhorn Helm repository on your management workstation or control-plane node:
helm repo add longhorn https://charts.longhorn.io
helm repo update
Next, create a custom values file named longhorn-values.yaml to tune replication parameters for your cluster. For two-node or three-node clusters, setting default replicas to 2 or 3 ensures balanced redundancy:
cat << 'EOF' > longhorn-values.yaml
defaultSettings:
defaultDataPath: "/var/lib/longhorn"
defaultReplicaCount: 2
storageOverProvisioningPercentage: 200
storageMinimalAvailablePercentage: 15
guaranteedInstanceManagerCPU: 10
replicaAutoBalance: "least-effort"
persistence:
defaultClass: true
defaultClassReplicaCount: 2
reclaimPolicy: "Delete"
csi:
kubeletRootDir: "/var/lib/kubelet"
EOF
Install Longhorn into its dedicated longhorn-system namespace:
helm install longhorn longhorn/longhorn \
--namespace longhorn-system \
--create-namespace \
-f longhorn-values.yaml
Monitor the rollout until all Longhorn CSI plugins, engine managers, and controllers reach Running status:
kubectl get pods -n longhorn-system -w
Once deployed, verify that Longhorn registered its default StorageClass:
kubectl get storageclass
You should see longhorn (default) alongside K3s’s built-in local-path provisioner.
Step 2: Exposing and Securing the Longhorn Dashboard
Longhorn provides an intuitive web interface for inspecting storage disks, real-time replica synchronization, volume throughput, and snapshot management. Because the dashboard has no built-in user authentication out of the box, exposing it directly to the internet is a severe vulnerability. We will secure it using Traefik Basic Authentication and Let’s Encrypt SSL.
First, generate an encrypted password hash using htpasswd:
# Install apache2-utils if htpasswd is not present
sudo apt install -y apache2-utils
# Create an auth secret with username 'admin'
htpasswd -nb admin "YourSecureClusterPasswordHere" > auth
Create the Kubernetes secret and Traefik Middleware inside the longhorn-system namespace:
cat << 'EOF' | kubectl apply -f -
apiVersion: v1
kind: Secret
metadata:
name: basic-auth-secret
namespace: longhorn-system
type: Opaque
data:
auth: $(cat auth | base64 -w 0)
---
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: longhorn-auth
namespace: longhorn-system
spec:
basicAuth:
secret: basic-auth-secret
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: longhorn-ingress
namespace: longhorn-system
annotations:
traefik.ingress.kubernetes.io/router.middlewares: longhorn-system-longhorn-auth@kubernetescrd
cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
ingressClassName: traefik
tls:
- hosts:
- longhorn.yourdomain.com
secretName: longhorn-dashboard-tls
rules:
- host: longhorn.yourdomain.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: longhorn-frontend
port:
number: 80
EOF
Navigating to https://longhorn.yourdomain.com now prompts for HTTP Basic Authentication before granting access to your cluster storage topology.
Step 3: Deploying a Stateful Workload with Replicated Storage
Let us test dynamic volume provisioning by deploying a production-ready PostgreSQL database with a 2-replica Persistent Volume Claim backed by Longhorn.
Create a dedicated namespace and deployment manifest:
cat << 'EOF' | kubectl apply -f -
apiVersion: v1
kind: Namespace
metadata:
name: storage-demo
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: postgres-data-pvc
namespace: storage-demo
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 20Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: postgres-demo
namespace: storage-demo
spec:
replicas: 1
selector:
matchLabels:
app: postgres-demo
template:
metadata:
labels:
app: postgres-demo
spec:
containers:
- name: postgres
image: postgres:16-alpine
env:
- name: POSTGRES_DB
value: "appdb"
- name: POSTGRES_USER
value: "dbadmin"
- name: POSTGRES_PASSWORD
value: "SuperSecretPassw0rd!"
- name: PGDATA
value: "/var/lib/postgresql/data/pgdata"
ports:
- containerPort: 5432
name: postgres
volumeMounts:
- name: db-storage
mountPath: /var/lib/postgresql/data
volumes:
- name: db-storage
persistentVolumeClaim:
claimName: postgres-data-pvc
EOF
Verify that Longhorn dynamically provisioned and bound the Persistent Volume:
kubectl get pvc,pv -n storage-demo
Inspect the Longhorn web dashboard. You will see postgres-data-pvc listed as Healthy, with two synchronous block replicas distributed across separate physical worker nodes.
Step 4: Automating S3 Snapshots and Disaster Recovery
Having synchronized replicas protects against single-node drive failures, but it does not protect against catastrophic datacenter outages, accidental table drops, or ransomware. Longhorn includes native S3 backup integration that takes incremental, deduplicated snapshots.
Create an AWS or MinIO S3 credentials secret in the longhorn-system namespace:
cat << 'EOF' | kubectl apply -f -
apiVersion: v1
kind: Secret
metadata:
name: s3-backup-secret
namespace: longhorn-system
type: Opaque
stringData:
AWS_ACCESS_KEY_ID: "YOUR_S3_ACCESS_KEY"
AWS_SECRET_ACCESS_KEY: "YOUR_S3_SECRET_KEY"
AWS_ENDPOINTS: "https://s3.eu-central-1.amazonaws.com"
EOF
In the Longhorn UI, navigate to Settings > General > Backup Target and configure:
- Backup Target:
s3://your-cluster-backups@eu-central-1/k3s-longhorn - Backup Target Credential Secret:
s3-backup-secret
Now, define an automated RecurringJob that takes a local snapshot every 6 hours and pushes an offsite backup to S3 every night at 02:00 UTC:
cat << 'EOF' | kubectl apply -f -
apiVersion: longhorn.io/v1beta2
kind: RecurringJob
metadata:
name: daily-s3-backup
namespace: longhorn-system
spec:
cron: "0 2 * * *"
task: "backup"
retain: 14
concurrency: 2
labels:
backup-policy: "daily"
EOF
Step 5: Testing Node Eviction and Failover Resilience
To verify that our distributed storage architecture truly prevents downtime when hardware fails, let us simulate a sudden node outage by cordoning and draining the active worker node hosting the PostgreSQL pod:
# Identify the node currently running the database
kubectl get pods -n storage-demo -o wide
# Cordon and drain the node
kubectl drain <worker-node-name> --ignore-daemonsets --delete-emptydir-data --force
Observe the Kubernetes scheduler and Longhorn in real time:
- The PostgreSQL pod on the drained node enters
Terminating. - Kubernetes reschedules the pod to the surviving worker node.
- The Longhorn CSI driver detaches the volume from the old node and attaches it to the surviving node in approximately 8 to 15 seconds.
- The PostgreSQL container mounts
/var/lib/postgresql/datafrom the surviving local replica and resumes serving database queries without data loss.
Once your maintenance simulation is complete, uncordon the node:
kubectl uncordon <worker-node-name>
Longhorn automatically detects that the node is online again, recalculates replica health, and synchronizes any missing block deltas in the background.
Troubleshooting Common Longhorn & K3s Issues
| Symptom | Root Cause | Resolution |
|---|---|---|
| Volume stuck in Attaching | Missing or inactive iscsid daemon on worker node. | Install open-iscsi on the host and run sudo systemctl restart iscsid. |
| Replica state Degraded | Cluster lacks enough distinct nodes or disks to satisfy replica count. | Adjust numberOfReplicas to match your active worker node count or add worker nodes. |
| Disk Out of Space | /var/lib/longhorn shared with OS root disk filled up. | Attach a secondary dedicated NVMe/SSD disk and configure it via Longhorn Node Disk settings. |
Summary & Next Steps
Deploying Longhorn on K3s elevates your Kubernetes infrastructure from an experimental single-node sandbox into a robust, fault-tolerant cluster. With synchronous block-level replication, intuitive web management, and automated offsite S3 backups, your databases and mission-critical containers can survive physical node outages with zero administrative panic.
To maximize Longhorn’s multi-node replication capabilities, your next architectural milestone is expanding your cluster topology by joining dedicated worker agent nodes to your control plane.
Hi, I’m Mark, the author of Clever IT Solutions: Mastering Technology for Success. I am passionate about empowering individuals to navigate the ever-changing world of information technology. With years of experience in the industry, I have honed my skills and knowledge to share with you. At Clever IT Solutions, we are dedicated to teaching you how to tackle any IT challenge, helping you stay ahead in today’s digital world. From troubleshooting common issues to mastering complex technologies, I am here to guide you every step of the way. Join me on this journey as we unlock the secrets to IT success.


