What causes I/O errors during a ZFS pool import?
On import, ZFS reads the vdev labels on each member, and the uberblock ring lives inside each label. Then it reads the Meta Object Set (MOS) that the best uberblock points to. An I/O error during this process means ZFS could not read at least one required metadata block.
- 1.Physical drive failure: One or more drives in the pool have failed or are returning read errors on sectors that contain ZFS metadata. If the failed drive is part of a raidz group and the group has exhausted its parity tolerance, import fails.
- 2.Controller or cable issues: If kernel logs show ATA errors or device resets preceding the ZFS I/O error, the problem may be at the transport layer, not the drive.
- 3.Damaged recent transaction groups: ZFS keeps a ring of uberblocks in each vdev label. On import it skips any uberblock that fails verification and uses the valid one with the highest TXG. A corrupt newest uberblock won't block import by itself. You only need to rewind to an older TXG when the pool won't import at that state.
- 4.Insufficient vdev members: raidz1 tolerates one failed drive. With two or more drives missing or faulted, the import fails. raidz2 tolerates two failed drives. Three or more faulted drives is past its redundancy, and the import fails.
What diagnostic steps should you run before taking action?
Gather information about the pool and drive health before attempting any import.
- 1.List available pools:
zpool import(no arguments) scans all connected drives for ZFS metadata and lists pools available for import. Note whether the pool shows as ONLINE, DEGRADED, FAULTED, or UNAVAIL. - 2.Read vdev labels:
zdb -l /dev/sdXreads the ZFS vdev label from a specific drive. This shows the pool name, pool GUID, vdev tree, and TXG number. Run on each drive to confirm all pool members are present and their labels are consistent. - 3.Check SMART data:
smartctl -a /dev/sdXfor each drive. Non-zero values in Reallocated_Sector_Ct, Current_Pending_Sector, or Offline_Uncorrectable indicate degraded media. - 4.Check kernel logs: Review
dmesgorjournalctl -kfor I/O errors, SATA link resets, or device timeouts. - 5.List uberblocks:
zdb -lu /dev/sdXreads the labels on one member and lists the uberblocks stored there with their TXG numbers. Run it on every member. If the pool won't import at its newest valid TXG, your rewind candidate is an earlier TXG from this list.
What are the safe ZFS pool import strategies?
If drives and cables are healthy and the issue is at the ZFS metadata level, these import strategies proceed from least invasive to most. Always attempt read-only import first.
- 1.Read-only import:
zpool import -o readonly=on poolnameimports the pool without writing metadata to the member drives; no TXGs are synced and no ZIL is replayed. Add-o cachefile=noneto prevent updatingzpool.cache. If any drive has suspected physical failure (clicking, slow response, I/O timeouts), image all drives with a hardware write-blocker first and import the cloned images instead. If ZFS can assemble the vdev tree, copy data to a separate destination immediately. - 2.TXG rollback: If read-only import fails because the latest TXG is corrupt, try importing at an earlier transaction group:
zpool import -T <txg-number> -o readonly=on poolname. Take valid candidates from the TXG numberszdb -lulists. At default settings, zfs_txg_timeout=5s is the longest a TXG runs, and it can close sooner. - 3.Force import (unclean shutdown only):
zpool import -f poolnametells ZFS to accept a pool that was not cleanly exported. ZFS replays the intent log (ZIL) and updates on-disk metadata. This is safe when all drives are healthy and the only issue is an unclean shutdown. If drives are failing, this writes to the drives and can overwrite recoverable metadata.
The order matters. Read-only import writes nothing to the pool's member drives. TXG rollback with readonly=on also writes nothing to the pool. Force import writes metadata updates. Each subsequent method is more invasive. If you skip straight to -f on a pool with failing drives, you may overwrite the blocks needed for recovery.
When can you resolve a ZFS pool import I/O error yourself?
Several common import failure scenarios have straightforward fixes that do not require professional recovery tools.
- 1.Device paths changed. Moving drives to a new controller, changing SATA ports, or booting a different OS can change
/dev/sd*assignments. ZFS identifies drives by GUID, not device path. Ifzpool importdoes not find the pool, tryzpool import -d /dev/disk/by-id/to scan by stable device identifiers. - 2.Unclean export after power loss. If the only issue is that the pool was not exported before shutdown,
zpool import -f poolnameis safe. ZFS replays the intent log. Confirm SMART data is clean on all drives before proceeding. - 3.Single drive failure in raidz2 or raidz3. If the I/O error is caused by one failed drive and the pool has sufficient parity margin, the pool may import in DEGRADED state. Replace the failed drive and resilver. Image all drives first if the data is irreplaceable.
- 4.Cable or controller port failure. If kernel logs show SATA link resets or device timeouts on a specific port, try a different cable or port.
When should you stop and image the drives?
Stop attempting imports and image every drive if any of the following conditions apply. Further import attempts on failing hardware risk making the data unrecoverable.
- 1.Multiple drives show I/O errors or SMART degradation. If more drives are faulted than the parity level allows (for raidz1, two or more; for raidz2, three or more), the pool has exceeded its redundancy. Further import attempts stress drives that are already failing.
- 2.Read-only import and TXG rollback both fail.
- 3.Previous repair attempts did not resolve the issue. If you've already run
zpool replaceorzpool scrubon the degraded pool and it didn't fix things, both of them wrote to the pool. A replace resilvers onto the new device, and a scrub repairs damage it finds. Image what remains before any further writes. - 4.Drives are making abnormal sounds. Clicking, grinding or repetitive seeking means a hardware fault. We find out which part failed on the bench, not from the sound. Power the drive down immediately. These drives need to be imaged with hardware that can manage weak heads and bad sectors at a level
ddrescuecannot reach.
For imaging, use ddrescue with a separate destination drive for each pool member. Work from the images for all subsequent recovery attempts. For drives with physical faults, professional NAS data recovery uses write-blocked connections and PC-3000/DeepSpar hardware to image drives that consumer tools cannot read, then reconstructs the pool offline from the images.
dRAID recovery reads the geometry fields off the vdev labels after imaging
dRAID is the declustered-parity topology added in OpenZFS 2.1. If your pool uses a dRAID vdev, we do the offline work with the same imaging discipline. dRAID permutes data, parity and spare capacity across every member instead of keeping them on fixed columns. You create the vdev with a suffix-tagged spec string:
# dRAID config string grammar: draid[parity]:[data]d:[children]c:[spares]s
draid2:4d:11c:1s
# parity = 2 redundancy level per redundancy group (1, 2, or 3)
# 4d = 4 data devices per redundancy group
# 11c = 11 total member devices in the vdev
# 1s = 1 distributed spare (sliced across all children, not a whole drive)The vdev label doesn't store the draid2:4d:11c:1s spec string. It stores the geometry as separate integer fields such as nparity, draid_ndata, draid_nspares and draid_ngroups. Read them off a surviving member with zdb -l /dev/sdX, where the vdev shows up as a draid type. A dRAID distributed spare is not a dedicated whole hot-spare drive. It is spare capacity sliced and interleaved across all children in the same permutation, so a sequential resilver writes into those distributed-spare slices rather than onto one separate disk. That fixed-width sequential rebuild reads from all children at once, which restores redundancy faster than a RAIDZ healing resilver. Fault tolerance is still the parity level: a draid2 vdev survives two lost members, and losing a third before the resilver finishes faults the vdev and takes the pool offline.
We image every member through a write-blocker, read the dRAID geometry fields off the labels, and reconstruct the pool on a forensic host. For enterprise TrueNAS handling of dRAID pools, see our TrueNAS server data recovery service.
Why does a pool refuse to import, or seem stuck after import, when every drive is healthy?
Two ZFS-specific conditions look like a failed import even when every pool member reads clean. A failed SLOG device makes zpool import refuse the pool until you use -m. A deduplication table (DDT) can take days to load on demand after import. Neither one prints an I/O error.
- 1.Failed SLOG device (ZIL loss): ZFS stages synchronous writes in the ZFS Intent Log (ZIL). When the pool has a dedicated Separate Intent Log (SLOG) device, the ZIL for sync writes lives on that SLOG until each record is committed into a transaction group (TXG) on the main pool. If the SLOG fails catastrophically following a sudden power loss, the in-flight synchronous writes staged on it cannot be replayed and are permanently lost. Only the unreplayed dirty sync writes are gone, not the whole pool. To import while discarding the missing log device, run
zpool import -m poolname. The-mflag accepts the loss of those unreplayed sync writes as the cost of bringing the pool online. Confirm the log vdev is the only missing device before using it. Forcing operations on drives that also show physical read errors is destructive because ZFS writes new dirty blocks over sectors you may still need. - 2.Slow DDT load after import: Enabling ZFS deduplication builds a Deduplication Table (DDT) proportional to the pool's unique block count. Holding the table in the Adaptive Replacement Cache (ARC) takes roughly 1 to 5 GB of RAM per 1 TB of deduplicated data. The DDT is stored on disk and read on demand, so a pool whose DDT does not fit in RAM still imports, but iXsystems documents that loading the table on demand can take days after an import or reboot, and every deduplicated write or deletion becomes a random disk read in the meantime. That slow-motion load reads to the operator like a failed import rather than the memory-sizing problem it actually is. Imaging or replacing healthy drives changes nothing. Because the DDT is not needed to read existing data, a strictly read-only import with
zpool import -o readonly=on poolnameis the recovery path on undersized hardware. Before import, estimate the DDT size withzdb -e -S poolname. It simulates deduplication and prints the DDT it would build.zpool status -Donly works on a pool that's already imported.
Both paths misdirect drive-level troubleshooting because the member disks are fine. See our TrueNAS data recovery service for how SLOG loss and DDT sizing factor into a full pool import.
Frequently Asked Questions
What does 'cannot import pool: I/O error' mean?
ZFS attempted to read metadata structures (vdev labels, uberblocks, or the Meta Object Set) from the pool's member drives and at least one read failed. It can mean a failed drive, a cable or controller problem, or too many missing vdev members for the pool's parity level. Run 'zpool import' without a pool name to see the pool's reported state, then check SMART data and kernel logs to narrow the cause.
Is 'zpool import -f' safe to run?
It depends on why the import failed. If the pool was not cleanly exported (unclean shutdown, power loss) but all drives are healthy, -f is safe. ZFS replays the intent log (ZIL) to recover synchronous writes that were acknowledged but not yet committed to a transaction group. Uncommitted asynchronous writes that existed only in RAM are lost. If drives are physically failing, -f forces ZFS to write to those drives, which can overwrite recoverable data. Check SMART data on every drive before using -f.
Can data be recovered from a FAULTED ZFS pool?
FAULTED means the pool has corrupted metadata or one or more faulted devices, and not enough replicas left to keep running. Recovery involves imaging each drive and reconstructing the pool offline. ZFS stores redundant copies of critical metadata (uberblocks, vdev labels) that forensic tools can use even when the live pool refuses to import.
My drives are all healthy, so why won't the pool import after a power loss?
If the pool had a dedicated SLOG device and it failed during the power loss, the synchronous writes staged in the ZFS Intent Log (ZIL) on that SLOG can't be replayed. Without -m, zpool import refuses the pool and tells you the missing devices can be skipped with '-m'. It doesn't print an I/O error. Only the unreplayed in-flight sync writes are lost, not the pool. Confirm the log vdev is the only missing device, then run 'zpool import -m poolname' to import while discarding the missing log. The -m flag accepts the loss of those unreplayed sync writes. Do not force this on drives that also show physical read errors, because ZFS will write new dirty blocks over sectors you may still need.
Why does my deduplicated pool take so long to come back after import?
ZFS deduplication builds a Deduplication Table (DDT) sized to the pool's unique block count. Holding it in ARC takes roughly 1 to 5 GB of RAM per 1 TB of deduplicated data. The DDT is stored on disk and read on demand, so the pool imports even when the table does not fit in RAM, but iXsystems documents that loading the table on demand can take days after an import or reboot, and every deduplicated write or deletion becomes a random disk read until it settles. That slow-motion load reads like a failed import. Imaging or swapping healthy drives does nothing. Because the DDT is not needed to read existing data, a strictly read-only import with 'zpool import -o readonly=on poolname' is the recovery path on undersized hardware. Before import, run 'zdb -e -S poolname'. It simulates deduplication and prints the DDT it would build, so you can estimate the table's size. 'zpool status -D' only works on a pool that's already imported.
Can you recover a TrueNAS dRAID pool that will not import?
dRAID is the declustered-parity topology added in OpenZFS 2.1. We put it through the same forensic process as any ZFS pool. We image every member through a write-blocker, then read the dRAID geometry off the vdev labels. The spec string is suffix-tagged: draid[parity]:[data]d:[children]c:[spares]s. For example, draid2:4d:11c:1s means parity level 2, 4 data devices per redundancy group, 11 total children and 1 distributed spare. The vdev label doesn't store that string. It stores the geometry as separate integer fields such as nparity, draid_ndata, draid_nspares and draid_ngroups, which 'zdb -l' prints. A dRAID distributed spare is spare capacity sliced across all children instead of one whole hot-spare drive. A sequential resilver writes into those slices. Fault tolerance is still the parity level.
Since 2008
Established
As Featured In
Related services
Related Recovery Services
Synology, QNAP, TrueNAS, and other NAS
Full RAID recovery service overview
Enterprise server recovery
FAULTED and UNAVAIL pool recovery
Synology NAS volume recovery
QuTS hero ZFS pool reconstruction
Transparent cost breakdown
ZFS pool import failing with I/O errors?
Free evaluation. Write-blocked drive imaging. Offline pool reconstruction with TXG history preserved. No data, no fee.