How to Deploy Scrutiny with Docker Compose for Real-Time Hard Drive S.M.A.R.T. Health Monitoring

In any production Linux server, homelab NAS, or virtualization hypervisor, storage drives are consumable components with finite lifespans. Mechanical hard disk drives (HDDs) suffer from motor wear, head friction, and platter degradation, while solid-state drives (SSDs and NVMe) wear out their NAND flash memory cells through program-erase (P/E) cycles. When a drive in a ZFS pool, hardware RAID array, or Unraid array fails silently without warning, the subsequent rebuild (resilvering) process stresses the remaining disks to their thermal and I/O limits, drastically increasing the risk of catastrophic multi-drive data loss.

Scrutiny Festplatten-Monitoring mit Docker Compose und SMART-Statusüberwachung
Real-time HDD and NVMe S.M.A.R.T. health monitoring with Scrutiny and Docker Compose

While the Self-Monitoring, Analysis, and Reporting Technology (S.M.A.R.T.) daemon (smartd) has been a Linux standard for decades, raw terminal outputs are notoriously cryptic. Manufacturers frequently report vendor-specific raw attributes, and standard email alerts are often either ignored or poorly configured. Scrutiny solves this challenge by transforming complex low-level smartctl metrics into an intuitive, real-time web dashboard paired with predictive failure analytics based on Backblaze’s vast drive reliability dataset.

In this guide, we will deploy a production-ready Scrutiny stack using Docker Compose. We will configure direct hardware device passthrough, set up automated multi-channel alerts (such as Discord and webhooks), explore the key S.M.A.R.T. indicators that signal impending mechanical and NAND failure, and review essential security hardening techniques.

Architecture: How Scrutiny Collects and Analyzes Drive Telemetry

Scrutiny follows a modular, client-server architecture consisting of three primary layers: the Collector daemon, the Web UI & API Service, and the InfluxDB Time-Series Database (bundled into an efficient all-in-one container or distributed across multiple hosts).

+------------------------------------------------------------------------------------+
|                               HOST HARDWARE & KERNEL                               |
|   +-------------------+  +-------------------+  +-------------------------------+  |
|   | /dev/sda (SATA)   |  | /dev/sdb (SAS)    |  | /dev/nvme0n1 (NVMe SSD)       |  |
|   +-------------------+  +-------------------+  +-------------------------------+  |
+------------------------------------------------------------------------------------+
                                          |
                        Direct Block Device Passthrough / udev
                                          v
+------------------------------------------------------------------------------------+
|                         SCRUTINY CONTAINER ENVIRONMENT                             |
|                                                                                    |
|  +------------------------------------------------------------------------------+  |
|  |                        SCRUTINY COLLECTOR DAEMON                             |  |
|  |    - Periodically executes smartctl / nvme-cli queries via cron schedule     |  |
|  |    - Normalizes raw vendor attributes (Seagate, WD, Crucial, Kioxia)         |  |
|  +------------------------------------------------------------------------------+  |
|                                         |                                          |
|                                         v (Internal Ingestion API)                 |
|  +------------------------------------------------------------------------------+  |
|  |                           CORE APPLICATION & API                             |  |
|  |    - Evaluates Backblaze drive failure thresholds and anomaly algorithms     |  |
|  |    - Dispatches instant alerts (Discord, Telegram, Email, Webhooks)          |  |
|  +------------------------------------------------------------------------------+  |
|                         |                                  |                       |
|                         v                                  v                       |
|            +-------------------------+        +-------------------------+          |
|            |    InfluxDB v2 / SQLite |        |    Responsive Web GUI   |          |
|            |    Time-Series Storage  |        |    Dashboard (Port 8080)|          |
|            +-------------------------+        +-------------------------+          |
+------------------------------------------------------------------------------------+
                                                              |
                                                    HTTP Browser Access
                                                              v
                                              [ System Administrator / SRE ]

The core architectural highlights include:

  • Attribute Normalization: Raw vendor numbers (such as Seagate’s notoriously high raw Seek Error Rate values) are parsed and translated into standardized health metrics.
  • Predictive Analytics: Compares your drive’s critical attribute changes against statistical failure patterns identified in enterprise datacenters.
  • Multi-Drive Topology Support: Accurately queries SATA, SAS behind LSI Host Bus Adapters (HBAs) in IT mode, and modern PCIe NVMe SSDs.
  • Centralized Hub-and-Spoke Deployment: Run the Web UI and Database on your primary server while deploying standalone lightweight collectors on remote virtualization nodes.

Step 1: Host Device Discovery & SMART Readiness Verification

Before launching the container, inspect your host’s block devices to identify all disks you wish to monitor and verify that S.M.A.R.T. monitoring is enabled at the hardware firmware level.

Run lsblk to enumerate physical block devices:

lsblk -d -o NAME,SIZE,MODEL,ROTA,TYPE

# Example Output:
# NAME    SIZE MODEL                     ROTA TYPE
# sda     16T  WDC WD161KRYZ-01A6NC0        1 disk
# sdb     16T  WDC WD161KRYZ-01A6NC0        1 disk
# sdc      2T  Samsung SSD 870 EVO 2TB      0 disk
# nvme0n1  1T  Samsung SSD 990 PRO 1TB      0 disk

Verify that smartmontools can successfully query your drives on the host:

# Install smartmontools if not present
sudo apt update && sudo apt install -y smartmontools

# Test SATA / SAS drive
sudo smartctl -i /dev/sda

# Test NVMe drive
sudo smartctl -i /dev/nvme0n1

If any drive displays SMART support is: Disabled, enable it permanently using:

sudo smartctl -s on /dev/sda

Step 2: Production Docker Compose Deployment

We will create a structured directory layout under /opt/scrutiny to hold configuration files, local database storage, and runtime logs.

sudo mkdir -p /opt/scrutiny/config
sudo mkdir -p /opt/scrutiny/influxdb

sudo chown -R 1000:1000 /opt/scrutiny
sudo chmod -R 755 /opt/scrutiny

Create /opt/scrutiny/docker-compose.yml:

services:
  scrutiny:
    image: ghcr.io/analogj/scrutiny:master-omnibus
    container_name: scrutiny
    restart: unless-stopped
    ports:
      - "8080:8080" # Web UI and Ingestion API
    environment:
      - PUID=0
      - PGID=0
      - TZ=UTC
      - SCRUTINY_WEB_INFLUXDB_INIT_MODE=setup
      - SCRUTINY_WEB_INFLUXDB_RETENTION_PERIOD=8760h # Retain 1 year of metrics
    volumes:
      - /opt/scrutiny/config:/opt/scrutiny/config
      - /opt/scrutiny/influxdb:/opt/scrutiny/influxdb
      - /run/udev:/run/udev:ro # Provides disk serial and hardware metadata
    # Access raw storage block devices:
    # Option 1 (Recommended for dynamic pools): Privileged mode
    privileged: true
    # Option 2 (Explicit granular device passthrough):
    # devices:
    #   - "/dev/sda:/dev/sda"
    #   - "/dev/sdb:/dev/sdb"
    #   - "/dev/nvme0n1:/dev/nvme0n1"
    #   - "/dev/nvme0:/dev/nvme0"
    cap_add:
      - SYS_RAWIO
      - SYS_ADMIN
    logging:
      driver: "json-file"
      options:
        max-size: "10m"
        max-file: "3"

Security Note on Privileged Mode: While running containers with privileged: true should generally be minimized, querying raw SCSI/ATA registers via smartctl and scanning dynamic disk controllers often requires CAP_SYS_RAWIO. If you prefer strict least-privilege containment, uncomment the explicit devices mapping list above and disable privileged: true.

Launch the stack and monitor container startup:

cd /opt/scrutiny
docker compose up -d

# Inspect startup logs
docker compose logs -f scrutiny

Step 3: Configuring Scrutiny Core & Collector Schedules

Scrutiny provides an extensive YAML configuration file for overriding default polling frequencies, drive filters, and threshold rules. Create or modify /opt/scrutiny/config/scrutiny.yaml:

version: 1

web:
  listen:
    port: 8080
    host: "0.0.0.0"

# Collector Execution Settings
collector:
  cron:
    # Run SMART health checks every 2 hours (default is every 15 minutes)
    schedule: "0 */2 * * *"

# Storage Device Filtering & Overrides
devices:
  # Automatically filter out transient virtual loop devices and docker virtual disks
  filter:
    exclude:
      - "/dev/loop.*"
      - "/dev/dm-.*"
      - "/dev/ram.*"

To run an immediate manual collection without waiting for the next cron interval, execute:

docker compose exec scrutiny /opt/scrutiny/bin/scrutiny-collector-metrics run

Open http://<server-ip>:8080 in your browser. You will see a modern dashboard listing every physical HDD and NVMe drive, complete with serial numbers, firmware revisions, current operating temperatures, power-on hours, and overall health status badges.

Step 4: Configuring Automated Multi-Channel Alert Notifications

A monitoring dashboard is ineffective if you are not proactively alerted when a drive encounters unrecoverable sector errors. Scrutiny integrates with Shoutrrr, providing seamless native notifications to Discord, Telegram, Pushover, Gotify, Slack, and standard SMTP email servers.

Add notification endpoints to your /opt/scrutiny/config/scrutiny.yaml file under the notify key:

notify:
  urls:
    # Discord Webhook integration
    - "discord://webhook_token@webhook_id"

    # Telegram Bot integration (Format: telegram://token@telegram?channels=channel_id)
    # - "telegram://123456789:ABCdefGhIJKlmNoPQRsTUVwxyZ@telegram?channels=-1001234567890"

    # Self-Hosted Gotify Server
    # - "gotify://gotify.yourdomain.internal/app_token?priority=8"

    # Standard Authenticated SMTP Email
    # - "smtp://user:password@smtp.mailprovider.com:587/?from=alerts@yourdomain.com&to=sysadmin@yourdomain.com"

  # Alert Trigger Criteria
  threshold:
    # Alert on any failure of critical SMART attributes
    failure: true
    # Alert if a metric is flagged as warning
    warning: true
    # Alert if drive temperature exceeds maximum threshold (in Celsius)
    temperature:
      critical: 55
      warning: 48

Restart the container to apply notification settings and test alert delivery:

docker compose restart scrutiny

# Trigger a test alert to verify webhook delivery
docker compose exec scrutiny /opt/scrutiny/bin/scrutiny-collector-metrics test-notification

A formatted alert message will immediately arrive in your configured channel confirming that notification delivery is functional.

Step 5: Decoding Critical S.M.A.R.T. Metrics & Failure Indicators

Not all S.M.A.R.T. attributes carry equal significance. While metrics like Power Cycle Count (Attribute 12) or Head Flying Hours (Attribute 240) are merely informational, Backblaze’s analysis of over 250,000 production drives demonstrates that a tiny subset of attributes correlates directly with imminent mechanical failure.

Critical Mechanical HDD Attributes (SATA / SAS)

ID Attribute Name Criticality Description & Failure Impact
05 Reallocated Sectors Count CRITICAL Raw count of damaged sectors moved to the spare reserve area. Any non-zero raw value indicates physical platter degradation. Replace immediately if increasing.
187 Reported Uncorrectable Errors CRITICAL Number of read/write operations that could not be recovered using hardware ECC. High correlation with drive failure within 60 days.
188 Command Timeout HIGH Number of aborted operations due to drive communication timeouts. Often points to dying controller logic or cable degradation.
197 Current Pending Sector Count CRITICAL Unstable sectors waiting to be remapped upon the next write operation. If writes fail, data stored in these sectors is permanently corrupted.
198 Offline Uncorrectable Sectors CRITICAL Quantity of uncorrectable sectors found during background autonomous self-tests. Direct indicator of bad disk sectors.

Critical NVMe Solid-State Metrics

Unlike legacy SATA drives, NVMe drives present standardized telemetry via the NVMe specification rather than legacy ATA attribute numbers:

  • Available Spare: Percentage (0% to 100%) of remaining reserved flash blocks available to replace degraded memory cells. If this drops below the manufacturer’s threshold (typically 10%), the drive enters read-only emergency mode.
  • Percentage Used: Vendor-estimated endurance consumed (can exceed 100% on heavily used drives). Serves as an odometer for warranty and write endurance.
  • Media and Data Integrity Errors: Raw count of unrecoverable data integrity errors encountered across the PCIe bus. Any value greater than 0 indicates imminent NAND controller or cell death.
  • Critical Warning Bits: A bitmask indicating temperature thresholds exceeded, backup power system degradation, or read-only fail-safe state.

Security Hardening & Production Best Practices

  1. Reverse Proxy Authentication: Scrutiny’s web interface does not feature built-in user login or password authentication by default. Never expose port 8080 directly to the public internet. Use a reverse proxy (such as Caddy, Nginx, or Traefik) configured with HTTP Basic Authentication, Authelia, or Authentik SSO.
  2. Dedicated Collector User: For multi-host setups, run only the lightweight collector binary on edge nodes, pushing JSON payloads over TLS to a centralized Scrutiny web hub without exposing raw host storage ports.
  3. Limit Polling Frequency: Polling mechanical drives every 5 minutes prevents them from entering low-power standby spin-down states. Set your cron schedule to once every 2 to 6 hours (0 */4 * * *) to prolong bearing life and conserve power.
  4. Backup Configuration Data: Include /opt/scrutiny/config and the underlying SQLite/InfluxDB directories in your automated 3-2-1 backup strategy (e.g., via Restic or Kopia) to preserve historical degradation charts.

Troubleshooting: Common Scrutiny & S.M.A.R.T. Issues

1. Error: “Cannot open device /dev/sdX: Permission denied”

Symptom: Scrutiny logs show failure to query drives, displaying smartctl open device failed: Permission denied.

Root Cause: The container process is running without sufficient Linux capabilities to execute raw device I/O against the kernel block subsystem.

Solution: Ensure your docker-compose.yml includes privileged: true or explicitly defines the cap_add permissions:

# Verify device permissions on host
ls -l /dev/sda

# Test running smartctl inside the container namespace
docker compose exec scrutiny smartctl -x /dev/sda

2. Issue: SAS Drives Behind an LSI MegaRAID / HBA Controller Missing Attributes

Symptom: Scrutiny displays SAS/SATA drives connected to a PCIe RAID controller or LSI HBA as a single combined virtual disk or fails to detect individual drives.

Root Cause: The controller requires explicit passthrough drivers (e.g., MegaRAID SAT passthrough or 3ware controller flags).

Solution: In /opt/scrutiny/config/scrutiny.yaml, add device-specific controller type flags under the devices section:

devices:
  overrides:
    - device: "/dev/bus/0 -d megaraid,0"
      type: "sat"
    - device: "/dev/bus/0 -d megaraid,1"
      type: "sat"

3. Issue: High InfluxDB Memory Usage & Disk Pool Load

Symptom: The Scrutiny container consumes excessive memory (> 2 GB RAM) and InfluxDB writes cause constant disk activity.

Root Cause: Running the collector too frequently (default 15 minutes) on pools with dozens of drives generates millions of time-series datapoints.

Solution: Lengthen the collector cron schedule to 4 or 6 hours in scrutiny.yaml, and set a finite InfluxDB retention policy in your environment variables:

environment:
  - SCRUTINY_WEB_INFLUXDB_RETENTION_PERIOD=4380h # 6 months retention

Conclusion

Hard drive and SSD failures are inevitable, but sudden unexpected data loss is entirely preventable. By deploying Scrutiny with Docker Compose, you gain deep, continuous visibility into your storage infrastructure without wading through cryptic raw terminal reports.

With standardized S.M.A.R.T. health ratings, Backblaze predictive failure algorithms, and instant multi-channel webhook notifications, you can identify failing storage media and replace degraded drives long before an array collapse compromises your data.