How Do You Tell Whether a QuTS Hero Pool Still Has Its Data Integrity?
Three per-vdev counters & one file list. zpool status reports READ, WRITE, & CKSUM counts for every member of every vdev, & zpool status -v names the files whose blocks no longer reconstruct. We read all four out of a read-only import of the cloned images, not out of the live unit.
What the zpool status counters actually report
READ & WRITE are the plain ones: both count I/O errors reported by the device itself. A non-zero CKSUM count means the device returned data that failed its checksum, which is silent corruption rather than a refused read.
The permanent-errors list ends the argument. zpool status -v lists files with unrepairable errors. That data is gone. Those specific files come back from a backup, not from another scrub.
QNAP's own QuTS hero documentation tells you to "use the Storage & Snapshots utility" & warns that "Using ZFS CLI commands on your QuTS hero NAS is not supported and should be avoided." That warning's right about the live unit. So we take the counters from a read-only import of the clones instead.
| Telemetry field | What it actually tells us |
|---|
| READ | I/O errors the device itself reported. A climbing count marks a member that gets imaged before anything else in the pool is touched. |
| WRITE | The same class of device-reported I/O error on the write path, which is why we read this counter from a clone instead of from a pool that is still accepting writes. |
| CKSUM | The device returned data that failed its checksum: silent corruption. |
| Permanent errors file list | Files with unrepairable errors. That data is gone. Those specific files come back from a backup, & no amount of further reading changes it. |
ZFS self-healing needs surviving redundancy
ZFS catching corruption doesn't mean ZFS can fix it. ZFS repairs a bad block as long as there is at least one good copy of that block that matches the checksum.
That good copy comes from redundancy: RAIDZ parity, a mirrored copy, or an extra copy ZFS kept because of the copies property. A RAIDZ vdev that has already lost its redundancy to a failed member has no surviving parity to rebuild that block from. ZFS still catches the checksum mismatch. It shows up as an unrecoverable read instead of a silent repair. That's the kind of error the permanent-errors list reports.
No scrub & no resilver until every member is imaged
A scrub reads every allocated block in the pool & verifies every checksum. A resilver only looks at data ZFS knows is out of date, like the data for a replaced device.
The mechanics of what that read load does to a same-batch vdev are in the resilver-triggered cascading drive failure walkthrough above. The rule that follows from it is short: clone every member first, then scrub the clones.
Snapshot retention limits inside a QuTS hero pool
A QuTS hero snapshot is an object inside the pool, not a copy parked somewhere else. It is enumerated with zfs list -t snapshot against an imported pool, the same dependency a zvol has on this platform: the pool assembles first, then the objects inside it become addressable.
Snapshots cover one kind of damage & do nothing for another. If an administrator overwrote a dataset, a snapshot from before the damage is the restore path. It won't help with corrupt vdev labels, dead member drives, or a broken uberblock ring. We reconstruct those offline before any snapshot is readable at all.
Plan retention around that split. Redundancy inside the chassis buys availability & uptime rather than data protection, so a copy that outlives a pool which won't import has to live off that pool: a replication target, or a discrete offline backup. Adding more snapshots to the same pool doesn't change that.