My Synology SHR volume crashed. What is SHR, and what do I do first?
SHR is not proprietary hardware. It is a software stack: mdadm for the RAID geometry, LVM to join the size bands, and Btrfs or ext4 on top. It reassembles on any Linux workstation. Power the NAS down, do not click Repair, and do not move the drives to a new Synology. Recovery means imaging every member, then rebuilding the mdadm and LVM layers offline. All work happens at our Austin, TX lab. Free evaluation, no data, no recovery fee.
Synology SHR Hybrid RAID Data Recovery
Your Synology shows a Volume Crashed banner, or an SHR array will not assemble after a rebuild. Before you touch anything, power the unit down.
You ship us your drives and all work happens at our Austin, TX lab. We clone every member through a write-blocker with the PC-3000 Portable III, the PC-3000 Express, or a DeepSpar Disk Imager before we analyze anything. Synology Hybrid RAID runs on the stock Linux md driver. That means we reconstruct it from those clones on a Linux workstation, and we don't need a Synology chassis to do it. It's part of our Synology NAS data recovery work. The evaluation is free, and if we can't recover your data, you don't pay a recovery fee.

What Is Synology SHR Actually Made Of?
SHR is a nested stack of standard Linux layers with a thin Synology management overlay on top. There is no proprietary RAID silicon inside a DiskStation. Knowing which layer failed decides whether your case is a routine read-only reassembly or a hand reconstruction, and none of these layers is a black box.
- 1. Physical disks and partitions
- On most models every member is partitioned the same way: a DSM system partition mirrored across the drives (
md0), a swap partition, and one or more data partitions. On a mixed-capacity array DSM cuts the larger drives into more than one data partition, so their spare capacity doesn't go to waste. - 2. mdadm software RAID
- The data partitions are aggregated by the kernel's md driver, managed from userspace with the mdadm tool. Each drive carries an mdadm 1.2 superblock written 4096 bytes into the partition, which records the array UUID, the member order, the chunk size, and the RAID level. A healthy array assembles with
mdadm --assemble --readonly. - 3. LVM logical volume
- The Storage Pool is an LVM volume group, the Linux Logical Volume Manager that lets a NAS combine mismatched drive sizes. On a mixed-capacity SHR array LVM concatenates several mdadm arrays end to end into one continuous logical volume. It is activated on a workstation with
vgchange -ay. - 4. Btrfs or ext4 filesystem
- The filesystem is formatted on top of the LVM logical volume. Btrfs is copy-on-write. On a normal write, it puts the new version of a block somewhere else and updates a pointer. That leaves the older tree roots sitting on disk. A repair tool that writes in place can destroy those older, still-valid versions of your data.
A standard equal-capacity SHR-1 of three or more drives is a single RAID 5 set under LVM and a Btrfs or ext4 filesystem. A mixed-capacity SHR is several RAID sets stacked by LVM. Either way we recover it from the clones in the same order: read the mdadm metadata, bridge with LVM, mount the filesystem read-only. The hard cases are the ones where the LVM bridge or the Btrfs tree is damaged.
What Is the Difference Between SHR-1 and SHR-2?
The number after SHR is how many drives can fail before the data is gone. SHR-1 survives one failure; SHR-2 survives two. The difference is which RAID level mdadm uses to build the size bands underneath, and that changes how much margin you have when a member is missing or unreadable during recovery.
| Attribute | SHR-1 | SHR-2 |
|---|---|---|
| Fault tolerance | One drive | Two drives |
| Minimum drives | One band can mirror on two drives | Four drives |
| Underlying mdadm level | RAID 5 on three or more drives, RAID 1 on a two-drive band | RAID 6 with dual parity once enough drives are present |
| Recovery margin | One missing or unreadable member per band before a stripe is unreconstructable | Two missing or unreadable members per band before loss |
The practical consequence shows up during a crashed-array recovery. On SHR-1 a single unreadable sector on a single-parity stripe has no second parity to reconstruct it, so that stripe is lost. SHR-2 can lose a whole second member and still solve every stripe, which is why the dual-parity layout is worth the extra drive on anything that matters.
For RackStation and larger DiskStation units, our enterprise and rack-mount Synology recovery page walks through the same imaging-first workflow at that scale.
Why Are Mixed-Capacity SHR Arrays Harder to Reconstruct?
A mixed-capacity SHR volume is not one RAID array. It is several stacked on top of each other and joined by LVM, and each one has to be solved separately if the metadata is gone.
Take an SHR-1 of two 1TB drives and two 2TB drives. DSM carves a 1TB partition on all four drives and a second 1TB partition on the two larger drives. It builds the first band as a four-drive RAID 5 across the matched 1TB partitions, and the second band as a two-drive RAID 1 mirror across the leftover space on the 2TB drives.
LVM then takes both mdadm devices, marks them as physical volumes in one volume group, and concatenates them into a single logical volume. That is how SHR uses all 6TB of raw space to give you 4TB usable with one-drive fault tolerance instead of wasting the extra capacity the way classic RAID 5 would.
When the superblocks survive, assembly is automatic. When they don't, we solve disk order, chunk size, and offset on the clones once per band, not once for the volume. The LVM metadata records the exact byte offset where the second band attaches to the end of the first. If that metadata is damaged, the filesystem won't mount until we rebuild that LVM mapping, even if every band reconstructs perfectly. That's why an LVM-layer failure is an in-lab job rather than something to attempt on the live NAS.
Why Did My SHR Volume Crash During a Rebuild?
One cause is an SMR drive timing out mid-rebuild. A drive-managed SMR drive shows up to the host as an ordinary hard drive. Under sustained writes it stalls for tens of seconds, and the NAS reads that stall as a dead drive.
An SMR drive writes overlapping tracks, like shingles on a roof. Rewriting data means rewriting a whole band. The drive hides that penalty behind a small zone of conventional CMR space used as a fast write cache.
A rebuild is a sustained, sequential, multi-terabyte write that runs flat out for hours. When the CMR cache overflows, the drive stalls for tens of seconds. The Linux kernel reads that stall as a dead drive and ejects the member from the array.
On a single-fault-tolerant SHR-1 that already lost one drive, that ejection is the second failure, and the pool collapses into the Volume Crashed banner. The cruel part is that the ejected SMR drive is physically healthy. It only timed out on cache overflow.
The ejected drive still holds its data. We image the array at a throttled pace with the timeout thresholds raised, then assemble it offline from the clones. It cannot be fixed by clicking Repair again, which only re-triggers the same stall.
An SMR ejection is one of several routes into the red banner, and the imaging-first triage for a Synology volume in the DSM Volume Crashed state is the same whichever fault got there first, even though the recovery work diverges once the cause is known. Every ejected member still gets imaged on its own, the same per-drive hard drive imaging a single dead disk gets, run once per bay.
Why Did My SHR Volume Crash After an NVMe Cache Failure?
When a read-write NVMe cache drops off the PCIe bus, the uncommitted writes it was holding go down with it. Your HDD array stays healthy. The Btrfs filesystem on top of it is left incomplete, and DSM reports the volume as crashed. Models with M.2 slots like the DS920+, DS1520+, and DS1621+ let you provision an NVMe SSD as cache, and the failure mode depends on which mode you chose.
- Read-only cache
- A read-only cache sits outside the write path. Synology's own SSD Cache white paper describes the cache as filled by data that is copied from the disks as it is requested, and it is the read-write mode, not this one, that Synology documents as write-back. In Synology's DSM 5.2 white paper, a multi-SSD read-only cache ran as RAID 0, and a read-write cache ran as RAID 1 "to ensure data integrity in case one SSD fails."
- Read-write cache (dangerous failure)
- A read-write cache sits in the write path. It intercepts new writes as a dirty cache and acknowledges them before they commit to the HDDs. When the NVMe drive drops off the PCIe bus, the writes it had acknowledged but not yet flushed to the array are gone. The HDDs never received them, and the filesystem is now missing structure it believes it already wrote.
Here is the part that decides the recovery: the corruption lands on the Btrfs filesystem layer, not on the mdadm or LVM block layers underneath it. The mdadm arrays and the LVM volume group come through the dropout intact, because the data that vanished with the cache was Btrfs filesystem-level metadata and extent updates, not block-device parity.
Btrfs is copy-on-write and stamps every tree update with a monotonically increasing transaction id, its generation tracking. When a parent node on the HDDs points at a child block and expects a transaction id that was only ever written to the now-vanished cache, the generation check fails.
The kernel reports a parent transid verify failed condition and refuses to mount the filesystem rather than serve corrupt structure.
Recovery runs on two tracks at once, the mechanical HDDs and the failed NVMe drive, because the surviving data and the lost data live on different media:
- HDD track: Image every mechanical HDD member through a hardware write-blocker with the PC-3000 Portable III, the PC-3000 Express, or a DeepSpar Disk Imager, assemble each size band with
mdadm --assemble --readonly, activate the LVM volume group, and mount the Btrfs filesystem read-only. Where the tree will not mount, we work read-only withbtrfs-find-rootandbtrfs restoreagainst historical generation roots that predate the aborted transaction. - NVMe track, in parallel: We image the failed NVMe drive separately.
There's a hard limit. If the NVMe drive is physically unrecoverable, the dirty writes are gone and nothing can recreate them.
btrfs-find-root and btrfs restore can still traverse older trees and pull out the files those trees still reference, but they can't bring back data that went down with the cache. What was acknowledged to the cache and never flushed is the part that is lost.
You'll find two pieces of advice on forums that make this worse. The first is to pull the cache SSDs to force the volume to mount. Synology warns that a read-write cache normally holds new data that hasn't been synchronized to the HDDs yet. It also warns that removing the SSDs before you remove the cache in Storage Manager can crash the volume.
The second is btrfs check --repair. It writes to the live tree and overwrites the historical copy-on-write roots we reach back to in a read-only extraction. Don't do either one on a crashed read-write cache volume.
If yours is one of the M.2-equipped prosumer units, the DS920+ NVMe read/write cache dropout is the exact case to read next, down to shipping the failed cache SSD alongside the drives. A caching device dying in front of the array isn't Synology-only either: a failed TrueNAS SLOG loses the in-flight synchronous writes after a sudden power loss, which is why NAS recovery across every vendor images the cache device and the array as two separate jobs.
Read-only forensic diagnostic (run against sector clones, never live drives)
# READ-ONLY DIAGNOSTIC. For sector-by-sector clones only,
# not live or degraded drives. The block stack is intact;
# the filesystem is the layer that will not mount.
# Block layers assemble cleanly from the HDD clones
# (N is the data partition number; Synology's own examples show 3 or 5)
mdadm --assemble --readonly /dev/md127 /dev/sdaN /dev/sdbN /dev/sdcN /dev/sddN
vgchange -ay
# A read-only mount fails with a generation mismatch when the
# cache that held the newest transaction never reached the HDDs:
# parent transid verify failed on <block> wanted <N> found <N-x>
# Traverse older generation roots read-only. No check --repair,
# no recovery mount framed as safe. These read; they never write.
btrfs-find-root /dev/vg1/volume_1
btrfs restore -t <older_root_bytenr> /dev/vg1/volume_1 /target/recoverThere is no btrfs check --repair here and no in-place repair. These commands read older trees; they never overwrite one.
Why Is Clicking Repair on a Degraded SHR Array Dangerous?
On a degraded array, Repair is a high-stress parity rebuild. There are two ways it turns a recoverable situation into a lost one. The first is mechanical.
Repair forces every surviving drive to read every sector to recompute the missing parity. If a surviving drive already has weak heads or bad sectors, that sustained read load can finish it off, turning a logical recovery into a clean-bench mechanical job.
The second is statistical. Consumer drives carry a worst-case rating of one unrecoverable read error per 10^14 bits read, which works out to roughly 12.5TB. That's the manufacturer's worst-case specification, not a countdown. Real-world drives often do far better.
But a degraded SHR-1 rebuild of, say, four 12TB drives has to read about 36TB off the surviving members to reconstruct the missing one, and across that much data the probability of hitting one latent unreadable sector is real.
On a single-parity stripe there is no second parity to rebuild that sector, so mdadm logs the unreadable block and carries on, but the data in that stripe is gone. That's the math. On a large array holding data with no verified backup, a RAID 5 or SHR-1 rebuild isn't a routine button press.
We do it the same way whatever the failure mode is. First we image every member with ddrescue, the PC-3000 Portable III, or a DeepSpar Disk Imager. We give any marginal drive a conservative retry profile so imaging doesn't accelerate wear. Then we recompute parity against the clones, where a mistake costs nothing. The Repair button writes to the only copy you have.
Why Should I Not Move SHR Drives to a New Synology or Run mdadm --create?
Two pieces of common advice can destroy a recoverable SHR array: migrating the drives to a replacement NAS, and forcing the array online with mdadm --create. Both write to the metadata you need for reconstruction.
When you move SHR drives into a new Synology and it offers to Migrate, DSM rewrites the md0 system partition so it can install its own operating system.
When the forced import then fails, DSM commonly offers a fresh install, which finishes the job. The metadata that recovery depends on is small and irreplaceable, and these workflows are precisely what overwrites it.
The mdadm --create command is just as final. Forum threads recommend it to force a crashed array back online, but create writes a brand new superblock over the original and makes you supply the exact drive order, chunk size, and layout from memory. Get one parameter wrong and that wrong geometry is now what the superblock says. Then mdadm starts a resync after the create, and it writes parity across that wrong layout, right over the data.
The original superblock is gone. The non-destructive assembly command is mdadm --assemble --readonly, and we run it against clones, never your live, degraded drives.
How Do We Read an SHR Superblock Without Touching the Data?
We run mdadm --examine /dev/sdXN against the cloned members, never the originals. It reads the mdadm 1.2 superblock at the 4096-byte offset without writing a single byte. It reports the array UUID, the Device Role (the member's slot), the chunk size, the RAID level, the array state flag, the event count, and the update time.
We can't assemble anything correctly until we've read that geometry back. Across every member we confirm the array UUID matches. The event counts show which member fell out of sync first, because the lowest event count is the member that dropped first. We recover the member order and chunk size we need to assemble the array, and we still haven't written a single byte. On a mixed-capacity SHR we solve this once per size band, since each band carries its own mdadm superblock with its own member order and chunk size.
Three mdadm verbs sound similar and behave nothing alike:
mdadm --examine /dev/sdXN- Read-only superblock interrogation. It doesn't write anything to the device, and it's the first command we run on each cloned member.
mdadm --assemble --readonly- Activates the array in read-only mode once we've confirmed the geometry. It reuses the existing superblocks rather than writing new ones.
mdadm --create- Writes a brand new superblock over the original and is destructive. It doesn't belong in a recovery workflow, because it overwrites the original metadata, and that's the metadata we need to reconstruct the array.
Read-only forensic diagnostic (run against sector clones, never live drives)
# READ-ONLY DIAGNOSTIC. For sector-by-sector clones only,
# not live or degraded drives. --examine writes nothing.
# Read the mdadm 1.2 superblock on one cloned data partition
mdadm --examine /dev/sdaN
# Reports: Array UUID, Device Role, Chunk Size,
# Raid Level, Array State, and Events (event count).
# Compare the UUID and Events across every member before any assembly.There is no --create here and no in-place repair. This command reads the superblock; it never writes one.
What Is the Difference Between Storage Pool Degraded and Volume Crashed?
Storage Pool Degraded and Volume Crashed are two real status strings that Synology DSM Storage Manager shows you, and they describe two different points in the same cascade. Degraded means the pool has lost members but is still within its fault tolerance. On SHR-1, a second drive fault on top of the first is what crashes it.
- Storage Pool Degraded
- On SHR-1, a single member has dropped. The array is still online and your data is still readable, but the redundancy is used up. There's no parity left to survive a second fault. If the data is irreplaceable and you have no verified backup, power down, label every drive by bay number, and image every surviving member, one at a time, through a write-blocker before anything else. Don't click Repair. A rebuild forces every surviving drive to read every sector and can trigger the second fault you're trying to avoid.
- Volume Crashed
- On SHR-1, a second drive has dropped out. The array is offline and the pool is inaccessible. To recover it, we image every member and reconstruct the mdadm and LVM stack offline from the clones.
On SHR-1 the single-member drop that shows as Degraded is the first fault, and a second drop, such as an SMR ejection, is what flips the status to Volume Crashed. Reading the event count with mdadm --examine on the clones is how the lab works backward through that cascade to establish which member left first.
How Do We Read the LVM Layer When Every mdadm Band Assembles?
There is a failure signature that sends the recovery somewhere else entirely: every size band assembles cleanly from the clones, the array UUID matches across all the members, the event counts agree, and the storage pool still refuses to activate. When the RAID layer checks out on every band, the damage is not in the RAID layer. It sits one level up, in the LVM metadata that concatenates those bands into the single logical volume DSM presented to you.
LVM is not a black box either. It writes three plain, documented structures onto the front of each physical volume, and the pvck man page states where each one lives.
- PV label: label_header and pv_header
- Both live inside one 512-byte sector, usually the second sector of the device. This is the stamp that says this block device is an LVM physical volume. Wipe it and LVM does not recognize the assembled mdadm band as a physical volume at all, so the band simply is not there as far as the volume manager is concerned.
- mda_header at offset 4096
- A 512-byte sector at a 4096-byte offset into the device, with an optional second copy near the end of the device. It is the pointer to the metadata text. Damage it and LVM can see that the device is a physical volume but cannot find the volume group descriptors it belongs to.
- Metadata text area
- An area immediately following that first mda_header sector, holding the volume group descriptors and the extent map in readable text. Zero it and the mapping that stitches the bands into one continuous logical volume is gone.
Read-only LVM header interrogation (run against sector clones, never live drives)
# READ-ONLY DIAGNOSTIC. For sector-by-sector clones only.
# These print LVM on-disk structures; they write nothing to the device.
# Print label_header, pv_header, mda_header(s) and metadata text
pvck --dump headers /dev/md127
# Print or save the volume group metadata text itself
pvck --dump metadata /dev/md127
# Report what LVM can currently see
pvs
vgs
lvsThe man page describes pvck --dump headers as printing the label_header, pv_header, mda_header(s) and metadata text, and warning when any value is incorrect. That is the whole job: read the structures, flag the bad ones, change nothing.
We run these the same way we run mdadm --examine: against the clones. A dump of the headers from each imaged band tells us which of the three structures is intact, which is damaged, and whether the second mda_header copy near the end of a device survived when the first one did not.
Two LVM commands are commonly recommended on forums as the fix, and both of them write. pvck --repair writes the header sector or the metadata area back to the device. vgcfgrestore rewrites volume group metadata onto the member devices from a text backup file, and it needs that file to carry the exact original UUIDs and extent mappings. If the backup is stale or belongs to a different layout, what it writes is a wrong map over the top of the only surviving correct one.
Restoring volume group metadata is a forensic operation for clones, not a line to type into a crashed DiskStation. On an ordinary Linux system, LVM keeps text-format copies of that metadata under /etc/lvm/backup and /etc/lvm/archive, and those files are what vgcfgrestore normally reads from.
We couldn't find a primary Synology source that says DSM keeps those files on its system partition. So nobody should assume that backup file is sitting on a crashed DiskStation waiting to be used. The copy of the metadata we count on is the one on the disks themselves, read out of the clones with pvck --dump.
How Do You Recover a Synology SHR Volume?
We image every member through a hardware write-blocker, reassemble the mdadm and LVM stack from the clones, and extract Btrfs or ext4 offline. Your original drives are never modified. This is lab work, not a do-it-yourself procedure, because a single wrong write to a superblock or an LVM header ends the recovery.
- Free evaluation: We document the model, the DSM error state, the SHR-1 or SHR-2 layout, the drive capacities and whether they are mixed, the filesystem, and which members were dropped and when. Suspected SMR members are flagged here so imaging is throttled from the start.
- Write-blocked imaging: Each member is cloned with the PC-3000 Portable III, the PC-3000 Express, or a DeepSpar Disk Imager. Marginal drives get conservative retry profiles and head maps; suspected-SMR drives get raised timeout thresholds so a cache-flush stall is not misread as a dead drive.
- Geometry and band detection: We read the mdadm 1.2 superblock at offset 4096 on each cloned data partition to recover the member order, chunk size, and RAID level for every size band. On a mixed-capacity array this is solved once per band rather than once for the volume.
- Read-only assembly: Each band is assembled from the clones with
mdadm --assemble --readonly, never with create. For arrays whose superblocks are too damaged to assemble, the bands are reconstructed virtually from the imaged members using Data Extractor Express RAID Edition on the PC-3000 Express. - LVM activation or repair: We list the logical volumes with
lvscanand activate them withvgchange -ayagainst the cloned images, never against your original drives. If the LVM metadata bridging one band to the next is damaged, we rebuild it by hand on the clones so the bands concatenate in the right order before we read the filesystem. - Filesystem extraction: The filesystem is mounted read-only. If the Btrfs tree is damaged we work read-only with
btrfs-find-rootandbtrfs restoreagainst historical generation roots. We never runbtrfs check --repairor force a recovery mount, because copy-on-write means an in-place write destroys the older roots the extraction depends on. - Verification and delivery: Recovered data is copied to a target drive, verified against your priority file list, and shipped back. Working copies are securely purged on request.
Read-only forensic diagnostics (run against sector clones, never live drives)
# READ-ONLY DIAGNOSTIC. For sector-by-sector clones only,
# not live or degraded drives. Wrong assembly = data loss.
# Assemble one SHR size band read-only from cloned partitions
# (N is the data partition number; Synology's own examples show 3 or 5)
mdadm --assemble --readonly /dev/md127 /dev/sdaN /dev/sdbN /dev/sdcN /dev/sddN
# List the logical volumes and activate the volume group (cloned images only)
lvscan
vgchange -ay
# Mount the filesystem read-only once the LV is active
mount -o ro /dev/vg1/volume_1 /mnt/recoverThese are diagnostics, not a repair guide. There is no --create step and no in-place filesystem repair, because both overwrite the metadata recovery depends on.
The same RAID and filesystem logic applies to Synology's classic RAID levels. SHR adds the size-band concatenation. Everything else works like any RAID data recovery case we run. SHR is one slice of the wider NAS data recovery work we handle across every vendor.
What Does SHR Data Recovery Cost?
SHR recovery uses the same two line items as any Synology array: a per-member price based on each drive's physical and firmware condition, plus one array reconstruction fee for the mdadm, LVM, and filesystem work. If we can't recover usable data, there's no recovery fee under our no-fix-no-fee guarantee.
Per-Member Drive Pricing
Each member drive is priced against the same five-tier schedule used for individual hard drive data recovery. A four-bay SHR unit with one head-swap member and three logical-only members generates an individual line item for each evaluated drive, not a single opaque bundle.
- Low complexity
Simple Copy
Your drive works, you just need the data moved off it
Functional drive; data transfer to new media
Rush available: +$100
$100
3-5 business days
- Low complexity
File System Recovery
Your drive isn't recognized by your computer, but it's not making unusual sounds
File system corruption. Accessible with professional recovery software but not by the OS
Starting price; final depends on complexity
From $250
2-4 weeks
- Medium complexity
Firmware Repair
Your drive is completely inaccessible. It may be detected but shows the wrong size or won't respond
Firmware corruption: ROM, modules, or translator tables corrupted; requires PC-3000 terminal access
CMR drive: $600. SMR drive: $900.
$600–$900
3-6 weeks
- High complexity
Head Swap
Bench diagnosis found the read/write heads have to be replaced. Clicking can also come from firmware, the preamp, or the spindle
Head stack assembly failure. Transplanting heads from a matching donor drive on a clean bench
50% deposit required. CMR: $1,200-$1,500 + donor. SMR: $1,500 + donor.
50% deposit required
$1,200–$1,500
4-8 weeks
- High complexity
Surface / Platter Damage
Your drive was dropped, has visible damage, or a head crash scraped the platters
Platter scoring or contamination. Requires platter cleaning and head swap
50% deposit required. Donor parts are consumed in the repair. Most difficult recovery type.
50% deposit required
$2,000
4-8 weeks
Hardware Repair vs. Software Locks
Our "no data, no fee" policy applies to hardware recovery. We do not bill for unsuccessful physical repairs. If we replace a hard drive read/write head assembly or repair a liquid-damaged logic board to a bootable state, the hardware repair is complete and standard rates apply. If data remains inaccessible due to user-configured software locks, a forgotten passcode, or a remote wipe command, the physical repair is still billable. We cannot bypass user encryption or activation locks.
No data, no fee. Free evaluation and firm quote before any paid work. Full guarantee details. Head swap and surface damage require a 50% deposit because donor parts are consumed in the attempt.
- Rush fee
- +$100 rush fee to move to the front of the queue
- Donor drives
- Donor drives are matching drives used for parts. Typical donor cost: $50–$150 for common drives, $200–$400 for rare or high-capacity models. We source the cheapest compatible donor available.
- Target drive
- The destination drive we copy recovered data onto. You can supply your own, or we'll provide one. For larger capacities (8TB, 10TB, 16TB and above), target drives cost $400+ extra. All prices are plus applicable tax.
Sealed helium drives are on their own price list, $200–$5,000+. When a head swap or platter repair opens one, we refill it with helium. That adds $400–$800, and the donor has to be an exact match. Helium drive prices
Array Reconstruction Fee
The array reconstruction fee is $400-$800. It covers mdadm parameter detection per size band, LVM reconstruction and any hex-level bridge repair, virtual assembly from cloned images, and Btrfs or ext4 extraction. It is confirmed at the free evaluation alongside the per-member line items.
No data, no recovery fee. If we can't recover usable data from your SHR volume, there's no recovery fee under our no-fix-no-fee guarantee. There are no diagnostic fees. A rush fee of $100 moves a case to the front of the imaging queue. Optional return shipping is the only other potential cost on an unsuccessful case.
How Do I Reduce the Risk of an SHR Volume Crash?
Use CMR drives from your NAS vendor's compatibility list, not whatever SMR drive was cheapest, because the SMR timeout cascade is the failure mode that turns a healthy drive into a crashed array during a rebuild. If you run a large SHR-1, understand that a degraded rebuild on high-capacity consumer drives carries real URE risk, and consider SHR-2 on anything you cannot afford to lose so a second unreadable member does not end the array.
SHR also does not change the oldest rule in storage: a redundant array gives you hardware availability, not a backup. Ransomware, an accidental deletion, a controller fault, or a cascading failure across drives from the same manufacturing batch destroys every member at once. Keep discrete, offline backups, and verify them with a test restore to a different machine before you assume they protect you.
How Do You Recover an iSCSI LUN From a Crashed SHR Array?
A file-level iSCSI LUN on Synology isn't a physical partition you can carve straight off the member disks. It's a file on the SHR volume's Btrfs or ext4 filesystem, inside the hidden @iSCSI directory. Synology says that directory holds LUNs and their snapshots.
We can't reach that file until we've reconstructed the full SHR stack read-only from the clones. That's the mdadm size bands, bridged by LVM, with Btrfs or ext4 on top. To recover the LUN, we do the same SHR reassembly, then extract the container file and loop-mount it:
- Reconstruct the SHR stack read-only: We image every member, reconstruct each SHR mdadm size band read-only, and bridge the bands with LVM, exactly as for a non-iSCSI SHR crash. We can't touch the LUN file until we've assembled the host volume underneath it.
- Mount the host filesystem read-only: We mount the Btrfs or ext4 host filesystem read-only and open the hidden
@iSCSIdirectory that holds the LUN container. - Extract the container file: We copy the LUN container file off the host filesystem to a healthy target drive.
- Loop-mount the extracted file: We attach the extracted file as a block device with
losetup -rPso the kernel scans the partition table the iSCSI initiator wrote inside the container. That exposes the inner NTFS, VMFS, or ext4 filesystem, and then we can read the data.
Whether the LUN is thin-provisioned or thick-provisioned changes how the container has to be handled during extraction:
- Thin-provisioned LUN
- A thin-provisioned LUN file is sparse: only the allocated extents hold real data, and the rest is unwritten holes. A naive flat copy that fills those holes can inflate the file to its full theoretical size and exhaust the target drive. The extract step has to preserve sparseness so a 2TB thin LUN holding 300GB of real data does not balloon into a 2TB flat file.
- Thick-provisioned LUN
- The container is fully allocated up front, so the file already occupies its declared size on the host filesystem. There are no holes to preserve, and the extracted file matches the size the initiator saw.
When the Btrfs extents holding the LUN container file are themselves damaged on a degraded SHR, we work read-only with btrfs-find-root and btrfs restore against historical generation roots, and never run btrfs check --repair. If the container's own extents are marginal, we ddrescue-image the extracted container file to a healthy target before loop-mounting it.
Two forum claims send people down dead ends. One is that a LUN is a physical partition you can carve straight off the member disks. The other is that it can be recovered without reconstructing the SHR mdadm, LVM, and Btrfs or ext4 stack underneath it. A file-level LUN is a file inside a host filesystem. We find the container file through that filesystem's own metadata, so we reconstruct the host volume first.
Read-only forensic diagnostic (run against sector clones, never live drives)
# READ-ONLY DIAGNOSTIC. For sector-by-sector clones only,
# not live or degraded drives. The host SHR stack is already
# assembled read-only before any of this runs.
# Mount the reconstructed SHR host filesystem read-only
mount -o ro /dev/vg1/volume_1 /mnt/recover
# Locate the LUN container in the hidden @iSCSI directory
ls /mnt/recover/@iSCSI/
# Extract the LUN container file (preserve sparseness on a thin LUN,
# or ddrescue it first if the container extents are marginal)
cp /mnt/recover/@iSCSI/<lun-file> /target/lun.img
# Loop-mount the extracted container read-only and read the inner FS
losetup -fPr /target/lun.img
mount -o ro /dev/loop0p1 /mnt/lunThese are diagnostics, not a repair guide. No btrfs check --repair, no in-place writes.
What Happens When an Encrypted SHR Volume Crashes?
An encrypted SHR crash is a two-track problem, and the honest answer is that one track is geometry and the other is a key you either have or you do not. On track one, we rebuild the same mdadm and LVM block stack as any SHR, from the clones. On track two, we unlock the LUKS container that DSM 7.2 layers on top of it.
DSM 7.2 full volume encryption is LUKS running in aes-xts-plain64 mode, the Linux device-mapper dm-crypt target. On an encrypted SHR volume that target sits between the LVM logical volume and the filesystem, which adds a fourth layer to the mdadm, LVM, and Btrfs or ext4 stack. Counting from the disks up, the order is now mdadm size bands, the LVM logical volume, the LUKS container, then the filesystem that lives inside the container. We reconstruct three of the four layers from on-disk metadata, with no secret involved. The LUKS layer is the only one whose recovery depends on a key.
- 1. mdadm software RAID (geometry)
- The data partitions aggregate into one or more mdadm size bands, exactly as on an unencrypted SHR. Each band carries its 1.2 superblock and assembles read-only with
mdadm --assemble --readonly. No key is involved at this layer. - 2. LVM logical volume (geometry)
- LVM concatenates the bands into one logical volume and activates with
vgchange -ay. On an encrypted volume the logical volume holds ciphertext rather than a mountable filesystem, but the layer itself is still recovered from LVM metadata alone, with no key. - 3. LUKS aes-xts-plain64 container (key-dependent, the new layer)
- This is the device-mapper dm-crypt target that DSM 7.2 places on top of the logical volume. It is the only layer that no amount of geometry work can open. To unlock it, we need the recovery key file DSM exported. We base64-decode it and hand it to
cryptsetup luksOpenas a key file, which creates the plaintext device mapping the filesystem lives inside. Without the key, the layer below it is intact and the layer above it is unreachable. - 4. Btrfs or ext4 filesystem
- The filesystem sits inside the unlocked LUKS container. Once the container opens, the plaintext mapping behaves like any SHR volume. We extract the Btrfs or ext4 filesystem read-only with the same tooling we use on any SHR, including
btrfs-find-rootandbtrfs restoreagainst historical generation roots where the tree is damaged.
The key material that auto-mounts the volume is held by the DSM Encryption Key Vault, local on the unit or remote over KMIP.
We work the two tracks in order. First we reconstruct the block stack on the clones. Once the LVM logical volume is active, we decrypt it by running cryptsetup luksOpen --readonly against the activated logical volume with the decoded recovery key file. That command maps a plaintext device without ever writing back to the encrypted container, and the Btrfs or ext4 filesystem is then mounted read-only off that mapping. If you exported that recovery key when you enabled encryption, the container opens and the rest is a standard SHR extraction.
An AES-XTS master key can't be brute-forced. If the Key Vault is unrecoverable and no recovery key was ever exported, decryption is impossible even with a flawless mdadm, LVM, and Btrfs block stack underneath. No forensic tool and no lab technique gets around that.
Any lab claiming a proprietary decryptor that opens a LUKS volume without the key is describing something that does not exist.
For encrypted volumes outside SHR, including file-level shared-folder encryption, see our Synology encrypted volume recovery page.
Read-only forensic diagnostic (run against sector clones, never live drives)
# READ-ONLY DIAGNOSTIC. For sector-by-sector clones only,
# not live or degraded drives. The block stack reconstructs
# from geometry; the LUKS layer needs the recovery key.
# Track 1: assemble the mdadm band and activate LVM read-only
mdadm --assemble --readonly /dev/md127 /dev/sdaN /dev/sdbN /dev/sdcN /dev/sddN
vgchange -ay
# Track 2: unlock the LUKS container on the logical volume.
# Requires the recovery key file DSM exported, base64-decoded.
# There is no way past this step without it; an AES-XTS key is
# not brute-forceable.
base64 --decode volume_1.rkey > /target/decoded_key
cryptsetup luksOpen --readonly -S 1 -d /target/decoded_key /dev/vg1/volume_1 shr_crypt
# Mount the filesystem inside the unlocked container read-only
mount -o ro /dev/mapper/shr_crypt /mnt/recoverThere is no btrfs check --repair here and no in-place repair. The container is opened read-only and the filesystem is read, never rewritten, and no command on this list can recover the data if the recovery key is gone.
What Are the Most Common SHR Recovery Questions?
Is Synology SHR proprietary hardware?
Can I use mdadm --create to recover a crashed Synology volume?
Can SHR be recovered without a Synology NAS?
Why did my SHR volume crash during a rebuild?
What is the difference between SHR-1 and SHR-2?
Should I click Repair in Synology Storage Manager?
My mixed-capacity SHR array will not assemble. Is it harder to recover?
What does Synology SHR recovery cost?
Can data be recovered from a crashed Synology encrypted volume?
Data Recovery Standards & Verification
Our Austin lab operates on a transparency-first model. We use industry-standard recovery tools, including PC-3000 and DeepSpar, combined with strict environmental controls to maintain drive integrity. This approach allows us to serve clients nationwide with consistent technical standards.
Localized Clean Zone
Open-drive work is performed in a 0.02 micron ULPA-filtered laminar clean bench.
Transparent History
Serving clients nationwide via mail-in service since 2008. Our lead engineer holds PC-3000 and HEX Akademia certifications for hard drive firmware repair and mechanical recovery.
Media Coverage
Our repair work has been covered by The Wall Street Journal and Business Insider, with CBC News reporting on our pricing transparency. Louis Rossmann has testified in Right to Repair hearings in multiple states and founded the Repair Preservation Group.
Aligned Incentives
Our "No Data, No Charge" policy means we assume the risk of the recovery attempt, not the client.
Technical Oversight
Louis Rossmann
Our engineers review all lab protocols to maintain technical accuracy and honest service. Since 2008, his focus has been on clear technical communication and accurate diagnostics rather than sales-driven explanations.
We believe in showing the bench rather than just describing it. Open-drive work runs on a 0.02 micron ULPA-filtered laminar clean bench, and we filmed it.
See the particle counter test at the benchSince 2008
Established
As Featured In
Related services
Related Recovery Services
Volume Crashed, Storage Pool Degraded, SHR reconstruction, and Btrfs/EXT4 recovery for DiskStation and RackStation units.
DSM 7.2 LUKS aes-xts-plain64 volumes and eCryptfs shared folders, with the honest limits when the key is gone.
NAS recovery for brands including QNAP, Buffalo, Western Digital, and Asustor.
Hardware and software RAID array reconstruction for RAID 0, 1, 5, 6, and 10.
Synology showing a Volume Crashed banner?
Free evaluation. No data, no recovery fee. Ship your drives from anywhere in the U.S.