Skip to main contentSkip to navigation
Lab Operational Since: 17 Years, 9 Months, 17 DaysFacility Status: Fully Operational & Accepting New Cases

SSD Causing System Freeze?
What the Stalls & Blue Screens Mean

A failing SSD freezes the machine because the host storage driver sits waiting on a read the drive can't finish. As NAND cells degrade the controller escalates to slower error correction, the command outlives the driver's timeout, & the driver resets the device. The system is frozen for that whole wait.

This is the state before the drive dies. It still enumerates, the operating system still boots off it, & the files are still there; the machine just locks up for seconds at a time whenever something touches the disk. That is the best window you will get for a clean image, and it closes the moment the controller stops answering. Our SSD data recovery work happens in-house at our Austin, TX lab, mail-in from anywhere in the country.

No diagnostic fee. No data, no recovery fee. Call (512) 212-9111

Author01/12
Louis Rossmann
Written by
Louis Rossmann
Founder & Chief Technician
Updated 2026-08-18
Mechanism02/12

Why Does a Failing SSD Freeze the Whole Computer?

A host read never touches a NAND cell directly. It asks the controller's Flash Translation Layer for a mapping, & the controller then reads the cells behind it. When those cells need repeated re-reads at shifted voltages, the answer arrives late, & the operating system stalls until it gives up.

Modern TLC and QLC NAND depends on the controller's LDPC soft-decision engine. Charge leaks out of the cells over time & with wear, threshold voltages drift, and the fast hard-decision decode that works on a healthy drive starts failing. The controller answers by re-reading the same page at finely shifted reference voltages until it can build a good enough probability map to correct the errors. It still returns your data. It just takes far longer to do it.

Published measurements in the flash-reliability literature put a clean hard-decision page read in the tens of microseconds, and a fully escalated soft-decision read past a millisecond. Multiply that by a directory walk, a game launch, or a virus scan & the storage stack falls behind by whole seconds. That is your freeze.

  1. Cell voltages drift. Trapped charge leaks through the tunnel oxide with age, heat, & program-erase cycling, and the raw bit error rate on those pages climbs.
  2. Hard-decision decoding fails. The controller reads each cell as a flat one or zero, finds more errors than that path can correct, & can't return the page.
  3. The controller escalates to soft decision. It walks its read-retry tables, re-sensing the same cells at shifted reference voltages to feed the LDPC engine a probability map instead of a flat bit.
  4. Latency climbs by orders. A read that used to complete in microseconds now takes milliseconds, and a request that spans many pages multiplies that penalty.
  5. The host driver gives up. When a command outlives the storage stack's timeout window, the driver aborts it and issues a reset to the device. Your machine is frozen for the length of that wait.
  6. The system recovers, or the volume drops. If the reset succeeds, the mouse starts moving again and you carry on. If it does not, the volume goes offline and applications start throwing I/O errors.

None of that is visible from the desktop, which is why the symptom gets misread as a slow computer. The mapping layer doing the work is covered in what the Flash Translation Layer does, and the cell physics behind the retries is in how NAND flash cells store data.

Evidence03/12

Which Log Entries Confirm a Storage Timeout?

Your operating system writes down every one of these stalls. The entry you want is a storage driver complaining about a reset or a retried command at the same wall-clock minute the machine locked up, not an application crash log. If the timestamps line up across several freezes, the drive is the suspect.

Windows
Event Viewer, under Windows Logs and then System. Event ID 129 records that a reset to the device was issued, which is the driver giving up on a command. Event ID 153 records an I/O operation that was retried, and Event ID 154 records one that failed with a hardware error.
Linux
Read dmesg or the journal. The nvme driver logs an I/O timeout, then a controller abort, then a controller reset as it escalates. A repeating cycle of those three lines is the same event Windows records as ID 129.
macOS
Open the panic reports in Console. A panic that names the NVMe controller driver points at the storage path; a panic naming a graphics or networking kext does not, no matter how convincing the freeze felt.

Representative log identifiers, not a capture from one machine

Windows  Event Viewer -> Windows Logs -> System
         Event ID 129   reset to device issued
         Event ID 153   I/O operation retried
         Event ID 154   I/O operation failed, hardware error

Linux    dmesg / journalctl -k
         nvme: I/O timeout
         nvme: controller abort
         nvme: reset controller

macOS    Console -> Crash Reports
         kernel panic naming the NVMe controller driver

One reset entry after a bad shutdown means nothing. A cluster of them, growing week over week, on a drive that has been in service for years, is the drive telling you how much longer it intends to cooperate.

Not the drive04/12

Is This the Windows 11 24H2 Host Memory Buffer Bug?

If the machine blue-screens under sustained writes on a DRAM-less WD or SanDisk NVMe drive, suspect the operating system before the hardware. Windows 11 24H2 changed Host Memory Buffer negotiation & tried to claim up to 200MB of host RAM, past what some drive firmware was written to handle.

A DRAM-less NVMe SSD has no cache chip of its own, so it borrows a slice of system RAM over PCIe to hold its mapping table. That negotiation historically topped out at 64MB. The 24H2 update pushed it to as much as 200MB, and the firmware on a specific set of Western Digital and SanDisk drives wasn't written for an allocation that size. Under sustained I/O the firmware panicked and took the host down with it.

The drives named in the reporting are the WD_Black SN770 and SN770M, the WD Blue SN580 and SN5000, and the SanDisk Extreme M.2. What matters more than the list is what the bug did & didn't do: it produced blue screens, and it left the controller silicon and the user data intact. No physical failure, no permanent data loss.

If that is your only symptom on one of those drives, do not pay anyone for data recovery. Update the drive firmware through WD Dashboard. The other documented mitigation is the HmbAllocationPolicy setting, which controls how much host memory the storage driver is allowed to hand a drive. Firmware first; the registry change is the fallback if you cannot flash yet.

A SATA SSD rules itself out here. Host Memory Buffer arrived with NVMe and runs over PCIe DMA, so a SATA controller can't negotiate it & a drive sitting behind a USB bridge can't use it either. If your stalling drive is SATA, or lives in a USB enclosure, this bug is not your problem and the read-latency mechanism above is back on the table.

Rule out05/12

What Else Stalls a Machine Besides the Drive?

A freeze on its own does not identify the drive. A lab that tells you otherwise before looking at your logs is guessing at your expense, so work through the cheap causes first. Each of these produces the same user-visible symptom and none of them needs a recovery invoice.

  • 01
    Bad RAM. A memory fault produces freezes, blue screens, & file corruption that all look like storage failure. Run a memory test to completion before you condemn the drive; a single failing address is enough to hang a machine daily.
  • 02
    Storage driver and chipset software. A vendor RAID or NVMe driver that doesn't match the platform, or a chipset package left behind after a Windows feature update, stalls the same I/O path the drive uses. Reverting to the inbox driver is a five-minute test.
  • 03
    Heat. An M.2 drive wedged under a graphics card with no heatsink throttles hard, and the stutter arrives a few minutes into any sustained transfer rather than at random. The drive's own temperature reading tells you within one file copy.
  • 04
    Power delivery. A failing power supply, a marginal riser, or a SATA power lead that has worked loose drops the drive mid-transfer. The host logs that as a device reset, which is the same entry a genuinely sick drive writes, so check the physical connection before reading anything into the log.

Work down that list and one of two things happens. Either the freezes stop, which costs you an afternoon, or they don't, & now you have evidence pointing at the drive instead of a hunch.

Health data06/12

NVMe Health Log Identifier 02h and the Critical Warning Byte

NVMe drives publish their health through a dedicated log page, the SMART / Health Information log at Log Identifier 02h. Any NVMe-aware utility can read it, and it's the closest thing to a straight answer the drive will give you about itself. Two fields on that page matter for a stalling drive.

The Critical Warning byte carries one condition flag per bit. The one that speaks to this symptom is bit 2, which the drive sets when NVM subsystem reliability has degraded. That flag is a threshold event, though. It trips when the firmware decides things are bad enough, and a drive can be losing the read-latency fight for months before that byte reports a thing.

Media and Data Integrity Errors is the more useful counter, because it moves earlier. It counts unrecovered data integrity errors, the cases where the controller exhausted its retries and still could not hand back correct data. A number that was zero last month and is climbing now tells you where this is going, regardless of what the summary health percentage claims.

One caution on tooling. SATA SMART uses numbered attributes with vendor-specific meanings, a different numbering scheme from the NVMe log page, so do not map an attribute ID from one onto a field in the other. If your drive is SATA, read it as SATA. Our SMART errors page covers what those attributes are worth on their own.

Do this now07/12

What Should You Do Right Now?

Stop writing to the drive & get an image before anything else. Copy the files you cannot replace to other media in one pass, then shut the machine down cleanly. Every extra hour of normal use runs more reads across the same cells the controller is already struggling to decode.

  1. Copy what you cannot replace, first. Skip the full clone and the whole-volume backup image for now. Pull the irreplaceable folders to an external drive in one pass and let the rest wait.
  2. Shut down cleanly. Use the menu, wait for it, and hold the power button only when the machine is genuinely unresponsive.
  3. Do not run CHKDSK, defrag, or Optimize Drives. Each of those walks the whole volume and writes as it goes, on a drive that is already failing to read cleanly.
  4. Do not run consumer deep-scan recovery software. A deep scan is a full-surface read by another name, and it fixes nothing that is wrong with a drive whose files still open.
  5. Check your drive model against the 24H2 list. If it is one of the DRAM-less WD or SanDisk NVMe drives above, flash the firmware and re-test before you spend anything.
  6. Power it down and send it in if the data matters more than the drive. Diagnosis is free, the quote is firm before any work starts, and mail-in works from anywhere in the country.

A firmware update is the right move on a healthy drive with a known firmware bug. It is the wrong move on a drive throwing uncorrectable read errors, because a flash rewrites controller structures on hardware that has already proven it cannot read reliably. Image first, flash second.

Escalation08/12

What a Forced Power Cycle Does to the FTL Mapping Table

The freeze pushes people toward the power button, and the power button is what turns a stalling drive into a dead one. The controller updates its mapping table constantly while it services reads, handles wear leveling, & runs garbage collection. Cut power during one of those metadata writes and the page lands half programmed.

Enterprise drives survive that with onboard capacitor banks that keep the controller alive long enough to finish the flush. Most consumer drives don't carry them. On the next boot the controller hits uncorrectable errors in its own service area, can't tell which copy of the map is authoritative, and locks itself into a firmware panic: safe mode, ROM mode, or a device that reports 0 bytes.

DRAM-less NVMe drives take this worse than most. Their active mapping table lives in borrowed host RAM across the PCIe link, so a sudden power loss severs the link and wipes that cache before any of it reaches NAND. That is the same architecture the 24H2 bug leaned on, so the DRAM-less models show up in both failure paths.

Once the controller is in that state, the symptom has changed and so has the work. A drive reporting 0GB or the wrong capacity and a drive Windows wants to initialize both start here. An NVMe drive that stops appearing in BIOS and a drive that has gone dead outright are the far end of it.

If the volume comes back mounted but unreadable, that is a RAW file system, and the capacity reporting behind it is in why an SSD reports zero bytes.

Software limits09/12

How Consumer Deep-Scan Software Makes a Stalling Drive Worse

Logical recovery software lives at the mercy of the controller. It issues standard ATA and NVMe reads through the operating system storage stack, has no path to vendor-specific commands, & no access to raw NAND pages or the spare-area metadata that a mapping-table rebuild requires. When the controller is the thing struggling, software has nothing to work with.

What it does have is a scan button, and that scan is a full-surface read. It drives sustained requests across the pages the controller is already retrying at shifted voltages, for hours, on a drive whose whole problem is that reads are expensive. CHKDSK and Optimize Drives add writes on top of that.

Deletions and cleanup during a repair attempt bring TRIM into it. The host operating system issues the deallocate command, not the drive; the controller then unmaps those addresses and synthesizes zeros for any later read of them, while garbage collection erases the physical blocks on its own schedule. Nothing at the software layer brings back an address the controller has unmapped. That mechanism is documented in what TRIM does and why it destroys data.

There is a narrow case where software is the right answer. If the drive passes a memory test and a driver rollback, the logs are clean, & you deleted a folder by accident on a machine where TRIM never ran, scan away. That is not what a freezing machine is describing.

In the lab10/12

How We Image a Drive That Still Responds

A consumer clone tool inherits the operating system's timeout, so it hangs on the same slow read that hangs your desktop, and it hangs again on the next one. Lab imaging hardware sets the timeout per command instead of accepting whatever the host driver decided.

Problem:
An SSD that still enumerates but stalls the host on individual reads, with rising read-retry latency on degrading NAND.
Hardware:
PC-3000 SSD with the PC-3000 SSD Extended add-on, PC-3000 Portable III
Outcome:
A sector-level image taken under controlled per-command timeouts before the controller drops into a firmware panic, with the file system rebuilt from that image.

The first pass runs short timeouts on purpose. Everything that reads quickly comes off fast, which on a drive in this state is usually most of it, and the addresses that stall get logged and skipped rather than retried until the controller overheats or gives up. Only then do we go back over the slow regions with longer timeouts, working the hard areas last instead of spending the drive's remaining life on them first.

PC-3000 SSD talks to a live controller over SATA or NVMe, which is the situation a stalling drive gives us. NVMe imaging runs on the PC-3000 Portable III, which can step the PCIe link speed down and adapt lane width when a drive is unstable at full negotiation. What that hardware does, & what it doesn't, is written up on our PC-3000 reference page.

Support for controller-level work is stratified, and pretending otherwise is how people end up paying for an evaluation that was never going to go anywhere. Logical and mapping-table work is supported on Phison S11, E12, E16, and E21T, on Silicon Motion SM2258XT, SM2259XT, SM2263XT, and SM2269XT, on Marvell 88SS1074, and on Samsung SATA MEX and MGX. Phison E26, Samsung in-house NVMe controllers, InnoGrit, Maxio NVMe, and SK Hynix proprietary silicon are board-repair only.

All of it happens at 2410 San Antonio Street in Austin, Texas. Single location, no franchises, no outsourcing, & the same shop since 2008. The technician who evaluates your drive is the one who images it. Full NVMe recovery and SATA SSD recovery coverage sits on the service pages.

Pricing11/12

What Recovery Costs When the Drive Still Reads

A drive that still enumerates and still returns data is usually an imaging job, not a chip-level one. Moving data off a functional drive runs $200. File-system recovery on a drive the OS won't mount starts at From $250. Firmware-level work runs $600–$900 on SATA.

What the drive is doingWork requiredTier
Stalls and stutters, volume still mounts, files still openControlled-timeout imaging and a straight copy off the imageSimple Copy, $200
Drive enumerates at the right capacity, volume won't mount or reads RAWImaging plus file-system reconstruction from the imageFile System Recovery, From $250
Controller has dropped into safe mode, ROM mode, or 0 bytes after a hard power-offController-level work to get the drive presenting data again, then imagingFirmware Recovery, $600–$900 SATA / $900–$1,200 NVMe

A drive that still reads belongs in that top row. Waiting until a forced reboot corrupts the mapping table moves the same drive down two rows. Rush service adds +$100 rush fee to move to the front of the queue. There is no diagnostic fee, the quote is firm before work begins, & no data means no recovery fee. Every tier is published on our pricing page.

Faq12/12

Common Questions

Data Recovery Standards & Verification

Our Austin lab operates on a transparency-first model. We use industry-standard recovery tools, including PC-3000 and DeepSpar, combined with strict environmental controls to maintain drive integrity. This approach allows us to serve clients nationwide with consistent technical standards.

Open-drive work is performed in a ULPA-filtered laminar-flow bench, validated to 0.02 µm particle count, verified using TSI P-Trak instrumentation.

Transparent History

Serving clients nationwide via mail-in service since 2008. Our lead engineer holds PC-3000 and HEX Akademia certifications for hard drive firmware repair and mechanical recovery.

Media Coverage

Our repair work has been covered by The Wall Street Journal and Business Insider, with CBC News reporting on our pricing transparency. Louis Rossmann has testified in Right to Repair hearings in multiple states and founded the Repair Preservation Group.

Aligned Incentives

Our "No Data, No Charge" policy means we assume the risk of the recovery attempt, not the client.

We believe in proving standards rather than just stating them. We use TSI P-Trak instrumentation to verify that clean-air benchmarks are met before any drive is opened.

See our clean bench validation data and particle test video

Image it while it still reads.

A stalling SSD is the easiest state to recover from and the shortest one. Free evaluation, firm quote before any work, no data no recovery fee.

(512) 212-9111Mon-Fri 10am-6pm CT
No diagnostic fee
No data, no fee
4.9 stars, 1,837+ reviews