
The OpenZFS filesystem has long been revered as the gold standard in enterprise storage, combining copy-on-write integrity, pooled storage management, and near-instantaneous atomic snapshots. However, snapshots by themselves do not constitute a backup. A local snapshot stored on the same physical zpool will not protect your infrastructure against disk controller failure, catastrophic pool corruption, or site-wide disasters. True disaster recovery requires replicating those snapshots offsite to an independent host.
While native tools like zfs snapshot, zfs send, and zfs receive provide raw block-level replication capabilities, orchestrating them via ad-hoc bash scripts frequently leads to brittle edge cases: unpruned snapshots silently filling disk pools, broken replication streams due to deleted intermediate snapshots, and sluggish transfers across low-bandwidth WAN connections.
Sanoid and Syncoid solve this problem systematically. Developed specifically for OpenZFS environments, Sanoid acts as a policy-driven snapshot management and pruning daemon, while Syncoid operates as an intelligent, resumable wrapper around zfs send and zfs receive. Together, they allow administrators to define granular retention policies (hourly, daily, weekly, monthly) and execute atomic, incremental snapshot replication over encrypted SSH tunnels with zero guesswork.
Architectural Overview: Policy-Driven Snapshots & Block-Level Sync
The Sanoid ecosystem cleanly divides storage hygiene from transport orchestration:
+-----------------------------------------------------------------------------------+
| PRODUCTION HOST (Primary Server / Proxmox / TrueNAS) |
| ZFS Pool: "tank/production" |
+-----------------------------------------------------------------------------------+
| |
| +---------------------------------------------------------------------------+ |
| | SANOID ENGINE (systemd timer: every 15 minutes) | |
| | | |
| | - Evaluates /etc/sanoid/sanoid.conf templates | |
| | - Generates atomic snapshots: autosnap_2026-10-02_18:00:00_hourly | |
| | - Automatically prunes expired snapshots according to retention policy | |
| +---------------------------------------------------------------------------+ |
| | |
| v |
| +---------------------------------------------------------------------------+ |
| | SYNCOID REPLICATION ENGINE | |
| | | |
| | - Calculates incremental differences between common snapshot bookmarks | |
| | - Pipes compressed block stream through lzop / zstd buffer | |
| | - Encrypts transit data over SSH using dedicated keyrings | |
| +---------------------------------------------------------------------------+ |
+-----------------------------------------------------------------------------------+
|
| Incremental "zfs send | ssh zfs receive"
| (Zero overhead, resumable, bandwidth-capped)
v
+-----------------------------------------------------------------------------------+
| BACKUP HOST (Offsite Storage / Secondary Homelab Node) |
| ZFS Pool: "backup-pool/replicated" |
| |
| - Receives exact block-level replicas with identical checksums |
| - Holds read-only mirror of production virtual machine disks & datasets |
| - Ready for immediate disaster recovery failover or point-in-time rollbacks |
+-----------------------------------------------------------------------------------+
Prerequisites and System Packages
Both Sanoid and Syncoid are written in Perl and leverage high-performance Linux utilities such as lzop, zstd, mbuffer, and pv to accelerate stream transfers and monitor progress.
Install the required packages on both your source (production) and target (backup) hosts running Ubuntu 22.04/24.04 or Debian 12:
# Update repositories and install Sanoid and supporting transfer tools
sudo apt-get update
sudo apt-get install -y sanoid lzop mbuffer pv zfsutils-linux
Verify that your ZFS pools are online and healthy:
sudo zpool status
Step 1: Configuring Granular Snapshot Policies in Sanoid
Sanoid uses a declarative configuration file located at /etc/sanoid/sanoid.conf. Policies are defined using reusable templates that specify how many hourly, daily, weekly, monthly, and yearly snapshots to preserve.
Create the Sanoid configuration directory and master configuration file:
sudo mkdir -p /etc/sanoid
sudo nano /etc/sanoid/sanoid.conf
Insert the following production-hardened configuration:
# ==============================================================================
# /etc/sanoid/sanoid.conf - Automated ZFS Snapshot Policy
# ==============================================================================
# Production VM Disks and Container Datasets
[tank/production/vms]
use_template = production
recursive = yes
# Database Storage (High-frequency hourly snapshots, shorter monthly retention)
[tank/production/databases]
use_template = database
recursive = yes
# Static Archival Storage (Infrequent snapshots, long multi-year retention)
[tank/production/archives]
use_template = archival
recursive = yes
# ==============================================================================
# RETENTION TEMPLATES
# ==============================================================================
[template_production]
frequently = 0
hourly = 24
daily = 30
weekly = 8
monthly = 12
yearly = 1
autosnap = yes
autoprune = yes
[template_database]
frequently = 4
hourly = 48
daily = 14
weekly = 4
monthly = 6
yearly = 0
autosnap = yes
autoprune = yes
[template_archival]
frequently = 0
hourly = 0
daily = 7
weekly = 4
monthly = 24
yearly = 5
autosnap = yes
autoprune = yes
In this configuration:
frequently = 4: Generates 15-minute interval snapshots (retaining the last 4 intervals, spanning one hour).hourly = 24: Keeps 24 hourly snapshots spanning the past 24 hours.daily = 30: Preserves one daily snapshot per day for the last month.autoprune = yes: Automatically destroys expired snapshots when their retention duration elapses, preventing capacity exhaustion.recursive = yes: Propagates the policy downward to all child datasets and zvols within the hierarchy.
Step 2: Testing Sanoid and Enabling Systemd Timers
Before leaving Sanoid to run autonomously, execute a dry run to inspect the snapshot operations it intends to perform:
# Perform dry run (no changes made)
sudo sanoid --cron --debug --dry-run
If the plan matches expectations, execute the initial snapshot creation manually:
# Take initial snapshots and prune
sudo sanoid --cron --verbose
Verify that the newly minted snapshots exist on your datasets:
sudo zfs list -t snapshot -r tank/production
Now enable and start the pre-packaged systemd timer that triggers Sanoid every 15 minutes:
sudo systemctl enable --now sanoid.timer
sudo systemctl status sanoid.timer
Step 3: Setting Up Dedicated Non-Root Replication via SSH
Running replication over root SSH without key restrictions poses a severe security hazard. Instead, create a dedicated unprivileged user (zfsbackup) on both hosts and leverage OpenZFS’s fine-grained delegation features (zfs allow).
On the target (backup) host, create the backup user:
sudo useradd -m -s /bin/bash zfsbackup
sudo passwd -l zfsbackup # Lock password authentication (keys only)
Delegate minimum required ZFS permissions to the zfsbackup user on the destination pool:
sudo zfs allow -u zfsbackup create,mount,receive,rollback,destroy,snapshot,hold backup-pool/replicated
On the source (production) host, generate an SSH keypair for the root or replication user:
sudo ssh-keygen -t ed25519 -f /root/.ssh/id_syncoid_backup -N "" -C "syncoid-replication"
Copy the public key (/root/.ssh/id_syncoid_backup.pub) to the destination server’s /home/zfsbackup/.ssh/authorized_keys file. Test the connection non-interactively:
sudo ssh -i /root/.ssh/id_syncoid_backup zfsbackup@backup-server.local "zfs list"
Step 4: Executing Block-Level Replication with Syncoid
With snapshot policies active and SSH connectivity established, Syncoid handles the actual data synchronization. Syncoid inspects the source and target datasets, identifies the most recent common snapshot, and transmits only the changed blocks.
Execute an initial replication sync with compression and progress monitoring:
sudo syncoid --no-sync-snap \
--compress=lzop \
--sshkey=/root/.ssh/id_syncoid_backup \
--recursive \
tank/production/vms \
zfsbackup@backup-server.local:backup-pool/replicated/vms
Notice the key flags:
--no-sync-snap: Tells Syncoid to replicate existing Sanoid-managed snapshots rather than creating transientsyncoid_...sync snapshots.--compress=lzop: Pipes data through the fastlzopcompression engine to reduce network utilization while maintaining 300+ MB/s throughput on multi-gigabit connections.--recursive: Recursively synchronizes all child datasets and volumes.
On subsequent runs, Syncoid completes in seconds because it only streams the delta blocks generated since the previous execution.
Step 5: Automating Syncoid Replication with Systemd
To run replication automatically (e.g., hourly), create a dedicated systemd service and timer unit on the source host.
Create /etc/systemd/system/syncoid-replication.service:
[Unit]
Description=Automated Syncoid ZFS Offsite Replication
After=network.target
[Service]
Type=oneshot
ExecStart=/usr/sbin/syncoid --no-sync-snap --compress=lzop --sshkey=/root/.ssh/id_syncoid_backup --recursive tank/production/vms zfsbackup@backup-server.local:backup-pool/replicated/vms
StandardOutput=journal
StandardError=journal
Create the corresponding timer in /etc/systemd/system/syncoid-replication.timer:
[Unit]
Description=Hourly Syncoid ZFS Offsite Replication Trigger
[Timer]
OnCalendar=hourly
Persistent=true
[Install]
WantedBy=timers.target
Reload systemd, enable, and start the timer:
sudo systemctl daemon-reload
sudo systemctl enable --now syncoid-replication.timer
sudo systemctl status syncoid-replication.timer
Step 6: Disaster Recovery and Snapshot Restore Procedures
A backup is only as good as its proven recovery workflow. There are two primary restoration scenarios depending on whether you need a fast local rollback or a full remote disaster recovery.
Scenario A: Fast Local Rollback
If a software upgrade corrupts a database, roll back to an atomic local snapshot in seconds:
# 1. Stop dependent services (e.g., Docker or database daemon)
sudo systemctl stop docker
# 2. Rollback to the desired snapshot
sudo zfs rollback -r tank/production/databases@autosnap_2026-10-02_17:00:00_hourly
# 3. Restart services
sudo systemctl start docker
Scenario B: Full Remote Disaster Recovery
If your primary storage server suffers a hardware failure, you can pull the entire dataset hierarchy back from your backup node:
# Pull dataset back from the remote backup server onto a newly installed zpool
sudo syncoid --compress=lzop \
--sshkey=/root/.ssh/id_syncoid_backup \
--recursive \
zfsbackup@backup-server.local:backup-pool/replicated/vms \
tank/production/vms
Troubleshooting Common Failure Modes
1. “Destination dataset has diverged” / Cannot Receive Incremental Stream
Cause: The target dataset was mounted as read-write or modified locally on the backup host, making native incremental fast-forward impossible.
Fix: Set readonly=on on all destination replica datasets: zfs set readonly=on backup-pool/replicated/vms. When running Syncoid, use --force-delete if you intentionally wish to overwrite destination drift.
2. Out-of-Space Errors Due to Hidden Old Snapshots
Cause: If a dataset previously had manual snapshots taken outside of Sanoid’s naming convention (e.g., manual_backup_...), Sanoid’s autoprune will deliberately ignore them to avoid data loss.
Fix: List all snapshots ordered by creation time: zfs list -t snapshot -o name,creation,used and prune forgotten legacy snapshots manually with zfs destroy <snapshot_name>.
3. Transfer Bottlenecks on High-Bandwidth LAN
Cause: Default SSH ciphers and single-threaded encryption can bottleneck multi-gigabit replication transfers.
Fix: In Syncoid, pass custom SSH options prioritizing fast modern ciphers: --sshoptions="-o Ciphers=chacha20-poly1305@openssh.com" or use an internal wireguard connection to eliminate SSH CPU overhead.
Conclusion and Next Steps
By pairing Sanoid’s automated retention policies with Syncoid’s efficient incremental replication, you establish an automated, resilient 3-2-1 backup strategy for Linux storage. With atomic point-in-time rollbacks available locally and block-accurate mirrors synchronized offsite, you gain complete protection against ransomware, filesystem corruption, and physical hardware loss.
To further extend this architecture, consider coupling Syncoid with an offsite cloud ZFS target (such as rsync.net or an offsite secondary TrueNAS appliance) or instrumenting your systemd timers with Prometheus node-exporter textfile collectors to monitor backup freshness in Grafana.
Hi, I’m Mark, the author of Clever IT Solutions: Mastering Technology for Success. I am passionate about empowering individuals to navigate the ever-changing world of information technology. With years of experience in the industry, I have honed my skills and knowledge to share with you. At Clever IT Solutions, we are dedicated to teaching you how to tackle any IT challenge, helping you stay ahead in today’s digital world. From troubleshooting common issues to mastering complex technologies, I am here to guide you every step of the way. Join me on this journey as we unlock the secrets to IT success.


