In any production Linux server, homelab NAS, or virtualization hypervisor, storage drives are consumable components with finite lifespans. Mechanical hard disk drives (HDDs) suffer from motor wear, head friction, and platter degradation, while solid-state drives (SSDs and NVMe) wear out their NAND flash memory cells through program-erase (P/E) cycles. When a drive in a ZFS pool, hardware RAID array, or Unraid array fails silently without warning, the subsequent rebuild (resilvering) process stresses the remaining disks to their thermal and I/O limits, drastically increasing the risk of catastrophic multi-drive data loss.

While the Self-Monitoring, Analysis, and Reporting Technology (S.M.A.R.T.) daemon (smartd) has been a Linux standard for decades, raw terminal outputs are notoriously cryptic. Manufacturers frequently report vendor-specific raw attributes, and standard email alerts are often either ignored or poorly configured. Scrutiny solves this challenge by transforming complex low-level smartctl metrics into an intuitive, real-time web dashboard paired with predictive failure analytics based on Backblaze’s vast drive reliability dataset.
In this guide, we will deploy a production-ready Scrutiny stack using Docker Compose. We will configure direct hardware device passthrough, set up automated multi-channel alerts (such as Discord and webhooks), explore the key S.M.A.R.T. indicators that signal impending mechanical and NAND failure, and review essential security hardening techniques.
Architecture: How Scrutiny Collects and Analyzes Drive Telemetry
Scrutiny follows a modular, client-server architecture consisting of three primary layers: the Collector daemon, the Web UI & API Service, and the InfluxDB Time-Series Database (bundled into an efficient all-in-one container or distributed across multiple hosts).
+------------------------------------------------------------------------------------+
| HOST HARDWARE & KERNEL |
| +-------------------+ +-------------------+ +-------------------------------+ |
| | /dev/sda (SATA) | | /dev/sdb (SAS) | | /dev/nvme0n1 (NVMe SSD) | |
| +-------------------+ +-------------------+ +-------------------------------+ |
+------------------------------------------------------------------------------------+
|
Direct Block Device Passthrough / udev
v
+------------------------------------------------------------------------------------+
| SCRUTINY CONTAINER ENVIRONMENT |
| |
| +------------------------------------------------------------------------------+ |
| | SCRUTINY COLLECTOR DAEMON | |
| | - Periodically executes smartctl / nvme-cli queries via cron schedule | |
| | - Normalizes raw vendor attributes (Seagate, WD, Crucial, Kioxia) | |
| +------------------------------------------------------------------------------+ |
| | |
| v (Internal Ingestion API) |
| +------------------------------------------------------------------------------+ |
| | CORE APPLICATION & API | |
| | - Evaluates Backblaze drive failure thresholds and anomaly algorithms | |
| | - Dispatches instant alerts (Discord, Telegram, Email, Webhooks) | |
| +------------------------------------------------------------------------------+ |
| | | |
| v v |
| +-------------------------+ +-------------------------+ |
| | InfluxDB v2 / SQLite | | Responsive Web GUI | |
| | Time-Series Storage | | Dashboard (Port 8080)| |
| +-------------------------+ +-------------------------+ |
+------------------------------------------------------------------------------------+
|
HTTP Browser Access
v
[ System Administrator / SRE ]
The core architectural highlights include:
- Attribute Normalization: Raw vendor numbers (such as Seagate’s notoriously high raw Seek Error Rate values) are parsed and translated into standardized health metrics.
- Predictive Analytics: Compares your drive’s critical attribute changes against statistical failure patterns identified in enterprise datacenters.
- Multi-Drive Topology Support: Accurately queries SATA, SAS behind LSI Host Bus Adapters (HBAs) in IT mode, and modern PCIe NVMe SSDs.
- Centralized Hub-and-Spoke Deployment: Run the Web UI and Database on your primary server while deploying standalone lightweight collectors on remote virtualization nodes.
Step 1: Host Device Discovery & SMART Readiness Verification
Before launching the container, inspect your host’s block devices to identify all disks you wish to monitor and verify that S.M.A.R.T. monitoring is enabled at the hardware firmware level.
Run lsblk to enumerate physical block devices:
lsblk -d -o NAME,SIZE,MODEL,ROTA,TYPE
# Example Output:
# NAME SIZE MODEL ROTA TYPE
# sda 16T WDC WD161KRYZ-01A6NC0 1 disk
# sdb 16T WDC WD161KRYZ-01A6NC0 1 disk
# sdc 2T Samsung SSD 870 EVO 2TB 0 disk
# nvme0n1 1T Samsung SSD 990 PRO 1TB 0 disk
Verify that smartmontools can successfully query your drives on the host:
# Install smartmontools if not present
sudo apt update && sudo apt install -y smartmontools
# Test SATA / SAS drive
sudo smartctl -i /dev/sda
# Test NVMe drive
sudo smartctl -i /dev/nvme0n1
If any drive displays SMART support is: Disabled, enable it permanently using:
sudo smartctl -s on /dev/sda
Step 2: Production Docker Compose Deployment
We will create a structured directory layout under /opt/scrutiny to hold configuration files, local database storage, and runtime logs.
sudo mkdir -p /opt/scrutiny/config
sudo mkdir -p /opt/scrutiny/influxdb
sudo chown -R 1000:1000 /opt/scrutiny
sudo chmod -R 755 /opt/scrutiny
Create /opt/scrutiny/docker-compose.yml:
services:
scrutiny:
image: ghcr.io/analogj/scrutiny:master-omnibus
container_name: scrutiny
restart: unless-stopped
ports:
- "8080:8080" # Web UI and Ingestion API
environment:
- PUID=0
- PGID=0
- TZ=UTC
- SCRUTINY_WEB_INFLUXDB_INIT_MODE=setup
- SCRUTINY_WEB_INFLUXDB_RETENTION_PERIOD=8760h # Retain 1 year of metrics
volumes:
- /opt/scrutiny/config:/opt/scrutiny/config
- /opt/scrutiny/influxdb:/opt/scrutiny/influxdb
- /run/udev:/run/udev:ro # Provides disk serial and hardware metadata
# Access raw storage block devices:
# Option 1 (Recommended for dynamic pools): Privileged mode
privileged: true
# Option 2 (Explicit granular device passthrough):
# devices:
# - "/dev/sda:/dev/sda"
# - "/dev/sdb:/dev/sdb"
# - "/dev/nvme0n1:/dev/nvme0n1"
# - "/dev/nvme0:/dev/nvme0"
cap_add:
- SYS_RAWIO
- SYS_ADMIN
logging:
driver: "json-file"
options:
max-size: "10m"
max-file: "3"
Security Note on Privileged Mode: While running containers with privileged: true should generally be minimized, querying raw SCSI/ATA registers via smartctl and scanning dynamic disk controllers often requires CAP_SYS_RAWIO. If you prefer strict least-privilege containment, uncomment the explicit devices mapping list above and disable privileged: true.
Launch the stack and monitor container startup:
cd /opt/scrutiny
docker compose up -d
# Inspect startup logs
docker compose logs -f scrutiny
Step 3: Configuring Scrutiny Core & Collector Schedules
Scrutiny provides an extensive YAML configuration file for overriding default polling frequencies, drive filters, and threshold rules. Create or modify /opt/scrutiny/config/scrutiny.yaml:
version: 1
web:
listen:
port: 8080
host: "0.0.0.0"
# Collector Execution Settings
collector:
cron:
# Run SMART health checks every 2 hours (default is every 15 minutes)
schedule: "0 */2 * * *"
# Storage Device Filtering & Overrides
devices:
# Automatically filter out transient virtual loop devices and docker virtual disks
filter:
exclude:
- "/dev/loop.*"
- "/dev/dm-.*"
- "/dev/ram.*"
To run an immediate manual collection without waiting for the next cron interval, execute:
docker compose exec scrutiny /opt/scrutiny/bin/scrutiny-collector-metrics run
Open http://<server-ip>:8080 in your browser. You will see a modern dashboard listing every physical HDD and NVMe drive, complete with serial numbers, firmware revisions, current operating temperatures, power-on hours, and overall health status badges.
Step 4: Configuring Automated Multi-Channel Alert Notifications
A monitoring dashboard is ineffective if you are not proactively alerted when a drive encounters unrecoverable sector errors. Scrutiny integrates with Shoutrrr, providing seamless native notifications to Discord, Telegram, Pushover, Gotify, Slack, and standard SMTP email servers.
Add notification endpoints to your /opt/scrutiny/config/scrutiny.yaml file under the notify key:
notify:
urls:
# Discord Webhook integration
- "discord://webhook_token@webhook_id"
# Telegram Bot integration (Format: telegram://token@telegram?channels=channel_id)
# - "telegram://123456789:ABCdefGhIJKlmNoPQRsTUVwxyZ@telegram?channels=-1001234567890"
# Self-Hosted Gotify Server
# - "gotify://gotify.yourdomain.internal/app_token?priority=8"
# Standard Authenticated SMTP Email
# - "smtp://user:password@smtp.mailprovider.com:587/?from=alerts@yourdomain.com&to=sysadmin@yourdomain.com"
# Alert Trigger Criteria
threshold:
# Alert on any failure of critical SMART attributes
failure: true
# Alert if a metric is flagged as warning
warning: true
# Alert if drive temperature exceeds maximum threshold (in Celsius)
temperature:
critical: 55
warning: 48
Restart the container to apply notification settings and test alert delivery:
docker compose restart scrutiny
# Trigger a test alert to verify webhook delivery
docker compose exec scrutiny /opt/scrutiny/bin/scrutiny-collector-metrics test-notification
A formatted alert message will immediately arrive in your configured channel confirming that notification delivery is functional.
Step 5: Decoding Critical S.M.A.R.T. Metrics & Failure Indicators
Not all S.M.A.R.T. attributes carry equal significance. While metrics like Power Cycle Count (Attribute 12) or Head Flying Hours (Attribute 240) are merely informational, Backblaze’s analysis of over 250,000 production drives demonstrates that a tiny subset of attributes correlates directly with imminent mechanical failure.
Critical Mechanical HDD Attributes (SATA / SAS)
| ID | Attribute Name | Criticality | Description & Failure Impact |
|---|---|---|---|
| 05 | Reallocated Sectors Count | CRITICAL | Raw count of damaged sectors moved to the spare reserve area. Any non-zero raw value indicates physical platter degradation. Replace immediately if increasing. |
| 187 | Reported Uncorrectable Errors | CRITICAL | Number of read/write operations that could not be recovered using hardware ECC. High correlation with drive failure within 60 days. |
| 188 | Command Timeout | HIGH | Number of aborted operations due to drive communication timeouts. Often points to dying controller logic or cable degradation. |
| 197 | Current Pending Sector Count | CRITICAL | Unstable sectors waiting to be remapped upon the next write operation. If writes fail, data stored in these sectors is permanently corrupted. |
| 198 | Offline Uncorrectable Sectors | CRITICAL | Quantity of uncorrectable sectors found during background autonomous self-tests. Direct indicator of bad disk sectors. |
Critical NVMe Solid-State Metrics
Unlike legacy SATA drives, NVMe drives present standardized telemetry via the NVMe specification rather than legacy ATA attribute numbers:
- Available Spare: Percentage (0% to 100%) of remaining reserved flash blocks available to replace degraded memory cells. If this drops below the manufacturer’s threshold (typically 10%), the drive enters read-only emergency mode.
- Percentage Used: Vendor-estimated endurance consumed (can exceed 100% on heavily used drives). Serves as an odometer for warranty and write endurance.
- Media and Data Integrity Errors: Raw count of unrecoverable data integrity errors encountered across the PCIe bus. Any value greater than
0indicates imminent NAND controller or cell death. - Critical Warning Bits: A bitmask indicating temperature thresholds exceeded, backup power system degradation, or read-only fail-safe state.
Security Hardening & Production Best Practices
- Reverse Proxy Authentication: Scrutiny’s web interface does not feature built-in user login or password authentication by default. Never expose port
8080directly to the public internet. Use a reverse proxy (such as Caddy, Nginx, or Traefik) configured with HTTP Basic Authentication, Authelia, or Authentik SSO. - Dedicated Collector User: For multi-host setups, run only the lightweight
collectorbinary on edge nodes, pushing JSON payloads over TLS to a centralized Scrutiny web hub without exposing raw host storage ports. - Limit Polling Frequency: Polling mechanical drives every 5 minutes prevents them from entering low-power standby spin-down states. Set your cron schedule to once every 2 to 6 hours (
0 */4 * * *) to prolong bearing life and conserve power. - Backup Configuration Data: Include
/opt/scrutiny/configand the underlying SQLite/InfluxDB directories in your automated 3-2-1 backup strategy (e.g., via Restic or Kopia) to preserve historical degradation charts.
Troubleshooting: Common Scrutiny & S.M.A.R.T. Issues
1. Error: “Cannot open device /dev/sdX: Permission denied”
Symptom: Scrutiny logs show failure to query drives, displaying smartctl open device failed: Permission denied.
Root Cause: The container process is running without sufficient Linux capabilities to execute raw device I/O against the kernel block subsystem.
Solution: Ensure your docker-compose.yml includes privileged: true or explicitly defines the cap_add permissions:
# Verify device permissions on host
ls -l /dev/sda
# Test running smartctl inside the container namespace
docker compose exec scrutiny smartctl -x /dev/sda
2. Issue: SAS Drives Behind an LSI MegaRAID / HBA Controller Missing Attributes
Symptom: Scrutiny displays SAS/SATA drives connected to a PCIe RAID controller or LSI HBA as a single combined virtual disk or fails to detect individual drives.
Root Cause: The controller requires explicit passthrough drivers (e.g., MegaRAID SAT passthrough or 3ware controller flags).
Solution: In /opt/scrutiny/config/scrutiny.yaml, add device-specific controller type flags under the devices section:
devices:
overrides:
- device: "/dev/bus/0 -d megaraid,0"
type: "sat"
- device: "/dev/bus/0 -d megaraid,1"
type: "sat"
3. Issue: High InfluxDB Memory Usage & Disk Pool Load
Symptom: The Scrutiny container consumes excessive memory (> 2 GB RAM) and InfluxDB writes cause constant disk activity.
Root Cause: Running the collector too frequently (default 15 minutes) on pools with dozens of drives generates millions of time-series datapoints.
Solution: Lengthen the collector cron schedule to 4 or 6 hours in scrutiny.yaml, and set a finite InfluxDB retention policy in your environment variables:
environment:
- SCRUTINY_WEB_INFLUXDB_RETENTION_PERIOD=4380h # 6 months retention
Conclusion
Hard drive and SSD failures are inevitable, but sudden unexpected data loss is entirely preventable. By deploying Scrutiny with Docker Compose, you gain deep, continuous visibility into your storage infrastructure without wading through cryptic raw terminal reports.
With standardized S.M.A.R.T. health ratings, Backblaze predictive failure algorithms, and instant multi-channel webhook notifications, you can identify failing storage media and replace degraded drives long before an array collapse compromises your data.
Hi, I’m Mark, the author of Clever IT Solutions: Mastering Technology for Success. I am passionate about empowering individuals to navigate the ever-changing world of information technology. With years of experience in the industry, I have honed my skills and knowledge to share with you. At Clever IT Solutions, we are dedicated to teaching you how to tackle any IT challenge, helping you stay ahead in today’s digital world. From troubleshooting common issues to mastering complex technologies, I am here to guide you every step of the way. Join me on this journey as we unlock the secrets to IT success.


