What Do the ZFS Pool States Mean?
ZFS tracks pool health at the vdev level. The four states that drive a recovery decision are ONLINE, DEGRADED, FAULTED, and UNAVAIL. The pool state is the worst state among its top-level vdevs.
ONLINE
All vdevs are healthy. No errors detected. Normal operation. No action needed.
DEGRADED
One or more vdevs have lost a member but can still serve data using redundancy (raidz parity or mirror copies). The pool is operational but running without full fault tolerance.
FAULTED
A vdev has lost too many members to maintain data integrity. raidz1 tolerates one drive failure; losing two or more exceeds its parity. raidz2 tolerates two failures; losing three or more exceeds its parity. The pool cannot serve I/O. Do not write to the remaining drives.
UNAVAIL
ZFS cannot open the device at all. The drive might be disconnected, it might have failed, or its device path might have changed. Check physical connections and device paths before assuming hardware failure.
How Do You Safely Export and Import a ZFS Pool?
Exporting a ZFS pool flushes pending writes and marks the pool as cleanly closed. Importing reads the on-disk metadata and reconstructs the in-memory state. Both operations are safe when done correctly, but import with the wrong flags can overwrite recoverable metadata.
- 1.
zpool export poolnameflushes all pending transaction groups to disk and marks the pool as exported. This is the cleanest way to take a pool offline. - 2.
zpool import -o readonly=on poolnameimports the pool in read-only mode, preventing the OS from writing metadata updates. Don't run it on drives you suspect are physically failing. Clone every member drive sector by sector on dedicated imaging hardware first, and import only the clones. - 3.
zpool import -f poolnameforces the import of a pool that appears to be potentially active, such as one that was never exported. That's a read-write import. It writes to the drives and can push the on-disk state past a point you could have recovered from. - 4.If the pool will not import at all, do not use
-frepeatedly. Image the drives and work from copies.
What Are the Risks of zpool clear?
zpool clear resets error counters on a vdev. If the errors were transient (a loose cable), this brings the vdev back online. If the drive is failing, clearing errors masks the problem and allows further corruption.
- 1.ZFS tracks read errors, write errors, and checksum errors per device. Once I/O errors go past acceptable levels, the device is faulted to prevent further use.
- 2.
zpool clear poolnameresets these counters to zero. If the drive answers, ZFS puts it back ONLINE.
Rule: Only run zpool clear if you have identified and fixed the root cause (reseated a cable, replaced a controller, resolved a power issue). If the drive itself is failing (check SMART), replace it instead of clearing errors.
How Do ZFS Transaction Group Rollbacks Work?
ZFS uses copy-on-write for dataset data and metadata. Changed blocks get written to a new location on disk, and the old blocks stay put until the space is reclaimed. Transaction groups (TXGs) are batched commits that advance the pool to a new consistent state. If the latest TXG is corrupted, you can roll back to a previous one.
- 1.ZFS flushes a new TXG to disk every 5 seconds (default) or when the write buffer fills.
- 2.The uberblock at the top of the pool metadata tree points to the most recent valid TXG. ZFS stores a ring buffer of historical uberblocks, providing a history of recent TXGs.
- 3.
zpool import -T <txg> -o readonly=on poolnameimports the pool using a specific historical TXG instead of the most recent. This effectively rolls back the pool to an earlier consistent state. - 4.TXG rollback only works if the on-disk blocks for the older TXG have not been overwritten by subsequent writes. Copy-on-write preserves old blocks until the space is needed.
- 5.The uberblock ring lives in the vdev labels on each member device, so read it off a device rather than off the pool name. Use
zdb -l -u /dev/sdXto list the uberblocks held in that device's labels and their TXG numbers before attempting a rollback.zdb -e -u poolnamedisplays the current uberblock. The-eflag tells zdb to work on an exported pool, one that isn't in /etc/zfs/zpool.cache.
Does Deduplication Stop a ZFS Pool From Importing?
No. The Deduplication Table lives on disk and ZFS reads it on demand, so a pool whose table doesn't fit in RAM still imports. What you lose is speed. Every deduplicated write or deletion turns into a random disk read. iXsystems says loading the table on demand can take days after an import or reboot.
Read-only forensic import path
A read-only import prevents ZIL replay. It never writes to the pool, so nothing about the attempt commits state to the drives. Adding -N suppresses automatic dataset mounting, so the host does not start touching datasets before you are ready to extract.
zpool import -o readonly=on -N <pool-name>Clearing the dedup property affects only newly written blocks; every block already written stays in the DDT. Eliminating a DDT means getting rid of the deduplicated blocks themselves: import the pool, copy the data to a fresh dataset with dedup off, and destroy the original.
We image every member with ddrescue or PC-3000 Portable III hardware first, then import the images read-only on a host sized for the table. All of that work happens in-house at the Austin, TX lab, and there is no diagnostic fee for the evaluation. For the appliance-level treatment of DDT sizing and SLOG loss, see TrueNAS ZFS pool import recovery. If the import fails with device errors, that's a different failure, and the ZFS pool import I/O error page covers it.
How Do You Handle FAULTED and UNAVAIL Vdevs?
When a vdev is FAULTED or UNAVAIL, the decision to replace or stop depends on the pool's remaining redundancy, the value of the data, and whether the failure is a drive issue or a connection issue.
- 1.Check physical connections first. UNAVAIL means ZFS couldn't open the device, and a changed device path is one cause. Reseat the SATA cable or try another controller port before you assume the drive failed.
- 2.Check SMART data. If the drive reports reallocated sectors, pending sectors, or UNC errors, the hardware is failing.
- 3.If the pool is DEGRADED (still has margin),
zpool replace poolname old-dev new-devinitiates a resilver (ZFS term for rebuild). That resilver reads the surviving vdev members to reconstruct the replacement. - 4.If the pool is FAULTED (no remaining margin), do not attempt to replace. The pool cannot guarantee data integrity. Image every drive and attempt offline reconstruction.
For enterprise server data recovery, FAULTED ZFS pools on production systems require imaging before any repair attempt. ZFS's copy-on-write architecture means the historical TXG data is still on-disk; any write operation (including a replace or scrub) can overwrite those blocks.
Resilver risk on unstable drives: Running zpool replace on a pool where the remaining drives have marginal heads or firmware instability puts a sustained read load on the surviving members. If the data is irreplaceable, image all members with ddrescue before initiating any resilver. For pools with missing top-level vdevs, the OpenZFS tunable zfs_max_missing_tvds allows importing a pool in read-only mode even when vdevs are absent, enabling data extraction without a full resilver.
How Do raidz1, raidz2, and raidz3 Differ in Fault Tolerance?
ZFS raidz levels map to traditional RAID data recovery parity concepts: raidz1 is single-parity (like RAID 5), raidz2 is dual-parity (like RAID 6), and raidz3 is triple-parity (no traditional RAID equivalent).
raidz1
Tolerates 1 drive failure.
raidz2
Tolerates 2 drive failures. Resilver can complete even with one URE.
raidz3
Tolerates 3 drive failures.
ZFS has one advantage over traditional RAID during rebuilds: it only resilvers allocated blocks, not the entire drive. This reduces both the time window and the total bytes read, lowering URE risk.
What Happens If the ZFS Encryption Key Is Lost?
If the wrapping key for an encrypted ZFS dataset is permanently lost, the data is gone and we cannot recover it. The passphrase or key file unwraps a randomly generated master key stored on disk in wrapped form; with no wrapping key, that master key stays sealed.
ZFS native encryption operates per dataset, not per pool and not at the block device layer. The pool imports normally with no key present. Topology reads, vdev labels verify, unencrypted datasets mount, and zpool status reports a healthy pool. The encrypted datasets sit unmounted. So "the pool imported" tells an owner nothing about whether their files are reachable.
zfs list -o name,keystatus,encryptionrootThat listing is the real status check. Each independently encrypted dataset, or a hierarchy of datasets inheriting one key configuration, carries an encryptionroot. A locked dataset shows keystatus of unavailable. Only after zfs load-key succeeds does it flip to available and allow a mount.
Where the key lives
- 1.A randomly generated master key encrypts the file data. It is created when the dataset is created and it is stored on disk in wrapped form. It never leaves the pool.
- 2.The user-supplied passphrase, raw key file, or hex key is the wrapping key, and it is held by the operator rather than by the pool.
zfs load-keyuses it to unwrap the master key into memory. - 3.With
encryption=onthe default suite isaes-256-gcm. Passphrases derive the wrapping key through PBKDF2, and thepbkdf2itersproperty defaults to 350,000 iterations and cannot be set below 100,000. - 4.
zfs change-keyis not a way in. It re-wraps the existing master key with a new wrapping key, and it requires the dataset to already be unlocked with the current key loaded. - 5.While a dataset is unlocked, the key material resides in system memory. On a live machine that is still unlocked, dumping kernel memory to extract key material is theoretically possible. Once the machine crashes, reboots, or loses power, that material is gone and no imaging of the drives brings it back. This is not a service we offer; it is stated here so nobody powers off a still-unlocked box expecting to pick up where they left off.
Lost wrapping key means we cannot recover the data. Without the passphrase, key file, or exported key configuration, the wrapped master key cannot be unwrapped. Sector-level imaging of every member drive changes nothing, because the ciphertext is intact and the problem is the missing key.
OpenZFS has no automatic key escrow and there is no hardware controller backdoor, because the key is held by the host rather than by a controller. AES-256 is not brute-forced.
What a lab can still report
ZFS native encryption protects the data payloads and the file names, not the enclosing filesystem structure. With no key at all, an evaluation can honestly tell an owner what existed: dataset names and the filesystem hierarchy, used capacity along with quotas and reservations, snapshot names and their creation timestamps, and dataset properties such as compression algorithm, record size, cipher, and PBKDF2 iteration count.
File names and file contents inside an encrypted dataset stay obscured. That inventory can tell somebody which of several pools held the data they care about, or that the snapshot they were counting on was never taken.
TrueNAS uses OpenZFS native encryption, not anything proprietary. If a key is set to unlock automatically at boot, it lives in the TrueNAS configuration database on the boot pool. Losing the boot device without a backup of that configuration or a manually exported key file loses the wrapping key, and the data pool stays sealed even though every data drive is healthy.
If you still have the passphrase, the key file, or an exported key, an encrypted pool falls under the TrueNAS and FreeNAS data recovery we do in-house at the Austin, TX lab.
We run this evaluation in-house at the Austin, TX lab with no diagnostic fee, and if the key is gone we say so and charge nothing. No data, no recovery fee applies here in its most literal form. Where a key does exist and the obstacle is pool structure rather than cryptography, the work returns to imaging every member and reconstructing the pool offline from the images.
Frequently Asked Questions
Can data be recovered from a FAULTED ZFS pool?
FAULTED doesn't mean the data is gone. It means ZFS has decided it can't guarantee data integrity with the vdevs it has available right now. The data's still on the drives. Recovery involves imaging each drive with a write-blocker and reconstructing the pool offline. ZFS stores extensive metadata, including multiple copies of the uberblock and transaction group history, which professional tools can use to rebuild the pool state even when the live pool refuses to import.
What does UNAVAIL mean in ZFS?
UNAVAIL means ZFS cannot open the vdev at all. The drive may have failed, been disconnected, or the device path may have changed. A single UNAVAIL vdev in a mirror is tolerated as long as the other mirror member is ONLINE. Check 'zpool status' for the specific vdev and drive identifier.
Is it safe to use zpool clear on a degraded pool?
zpool clear resets the error counters on vdevs that ZFS has flagged. If the errors were transient (a loose cable, a temporary controller issue), clearing can bring the vdev back online. If the errors reflect a real hardware failure, clearing masks the problem and allows the pool to continue operating with a drive that is actively failing. Only use zpool clear if you have identified and resolved the root cause of the errors.
Why won't my ZFS pool import when every drive reports ONLINE?
If every member is ONLINE, the metadata verifies, and the pool is so slow it looks stuck instead of throwing an error, check whether deduplication was enabled. The Deduplication Table lives in on-disk ZAP objects and is read on demand rather than loaded at import, so the pool imports either way; what scales with the amount of deduplicated data is how much RAM it takes to keep that table cached in ARC. Below that, DDT entries page off disk and performance craters: iXsystems documents that loading the table on demand can take days after an import or reboot, and every deduplicated write or deletion becomes a random disk read. Keeping that table in ARC takes roughly 1 to 5 GB of RAM per TB of deduplicated data. Clearing the dedup property affects only newly written blocks. Because the DDT is not needed to read existing data, import cloned images read-only with zpool import -o readonly=on -N.
Can a data recovery lab decrypt a ZFS dataset if the encryption key is lost?
No. We cannot, and neither can anyone else. ZFS native encryption stores a randomly generated master key on disk in wrapped form, and the passphrase, raw key file, or hex key you supply is the wrapping key that zfs load-key uses to unwrap it. With the wrapping key permanently gone, the master key stays sealed. Imaging every member drive changes nothing, because the ciphertext is intact and the missing piece is the key. OpenZFS has no key escrow and there is no controller backdoor, since the key is held by the host. Note that the pool itself still imports without a key: encryption is per dataset, so the encrypted datasets show keystatus unavailable in 'zfs list -o name,keystatus,encryptionroot'. What an evaluation can still report from a locked dataset is dataset names and hierarchy, used capacity and quotas, snapshot names with creation timestamps, and dataset properties. File names and file contents stay obscured.
Since 2008
Established
As Featured In
Related services
Related Recovery Services
ZFS pool FAULTED or UNAVAIL?
Free evaluation. Write-blocked drive imaging. Offline pool reconstruction with TXG history preserved. No data, no fee.