Fix it now
0x00000124 is a fatal hardware error reported through the Windows Hardware Error Architecture. Parameter 1 says which error source raised it and parameter 2 is the address of the error record, which is the only thing that tells you what actually failed.
wevtutil qe System /c:40 /rd:true /f:text
- Read the System log around the crash. Hardware error entries logged before the bug check name the component far more usefully than the stop screen does, and Microsoft’s own resolution list starts with checking Event Viewer.
- Undo any overclock, including memory profiles and any board-level automatic tuning. Microsoft names over-clocking in the cause list for this code.
- Check that cooling is working: fans spinning, heatsink seated, vents clear, and the machine not thermally throttling under load.
- Get parameter 1 from the BugCheck event or the dump. 0x0 is a machine check exception, 0x3 an NMI, 0x4 an uncorrectable PCI Express error, 0x10 a device driver error source – and each sends you somewhere different.
- Update the system firmware and the chipset package from the machine or board vendor, then re-test.
- Run the hardware diagnostics the system manufacturer supplies, and Windows Memory Diagnostics on top of them.
If you have the dump open in WinDbg, !errrec against parameter 2 decodes the WHEA error record, which names the component far more precisely than anything on the stop screen.
If the machine is stable with the overclock removed and firmware current, stop here. The next section covers reading the error source and the codes that look like this one but are not.
Why it happens
The Windows Hardware Error Architecture is the path by which platform hardware reports errors to the operating system. Processors, memory controllers, PCI Express root ports and platform error sources all feed into it. Correctable errors are logged and life goes on. When an error arrives that cannot be corrected and cannot be contained, Windows bug checks with 0x00000124 and records what it was told.
That recording is the point. Parameter 2 is the address of a WHEA_ERROR_RECORD structure, and !errrec on it in the debugger produces the specific error: which processor bank, which memory address, which PCI Express device. Without it you are guessing which component to swap, which is why this code produces so much fruitless hardware replacement.
Parameter 1 identifies the error source and narrows the field before you even open the dump. 0x0 is a machine check exception and 0x1 a corrected one; 0x2 a corrected platform error; 0x3 a non-maskable interrupt; 0x4 an uncorrectable PCI Express error; 0x5 a generic hardware error; 0x6 an initialisation error; 0x7 a BOOT error; 0x10 a device driver error source; 0x11 and 0x12 Arm synchronous external abort and SError interrupt. A 0x4 sends you to an add-in card or a riser; a 0x0 sends you to the processor, memory or the path between them.
Microsoft’s cause list is heat, defective hardware, memory or a processor beginning to fail, and over-clocking – and then, explicitly, ‘less likely, but possible: a driver is causing the hardware to fail with this bug check’. That last line matters. It is not true that software cannot cause this, and a machine that only produces 0x124 under one workload with one device attached is worth investigating as a driver problem before it is stripped for parts.
The machine is overclocked or running outside stock settings
You have this one if Any memory profile, base clock change, voltage offset or board automatic tuning is enabled. Crashes cluster under load.
- Load firmware defaults and retest without any profile enabled. That includes memory profiles, which are overclocks whatever the marketing calls them.
- If the machine is stable at stock, reintroduce one setting at a time.
- Where an undervolt or a power limit has been applied for thermal reasons, remove that too while testing – instability is instability whichever direction the voltage moved.
- Confirm the firmware is current before you conclude the silicon is at fault.
Heat or a cooling failure
You have this one if Crashes follow sustained load, or the machine is in a warm enclosure, or fans have been quiet since the last time it was opened.
- Confirm every fan spins and the heatsink is seated with intact thermal contact.
- Clear dust from intakes and radiator fins, and check the machine is not recirculating its own exhaust.
- Watch temperatures under sustained load with the vendor’s own utility.
- Retest after the machine has been running cool for the duration of a full workload.
Failing memory or processor
You have this one if Parameter 1 of 0x0, the error record naming a memory address or a processor bank, and no overclock in play.
- Run Windows Memory Diagnostics with the extended test mix, and the manufacturer’s own diagnostics on top.
- Test one memory module at a time where more than one is fitted.
- Read the error record with
!errrecand take the component it names to the vendor. - On a machine under warranty, the decoded record is the evidence that gets a part replaced.
A PCI Express device or its driver
You have this one if Parameter 1 of 0x4, an uncorrectable PCI Express error, or 0x10, a device driver error source.
- Reseat the card and try a different slot. Risers and adapters are a common source.
- Update the device’s firmware and driver from its own vendor.
- Remove the card entirely and see whether the machine is stable without it.
- Where the source is 0x10, the reporting component is a driver, so treat it as a driver investigation rather than a component swap.
Full reference
Parameter 1: the error source types
| Value | Error source |
|---|---|
0x0 |
A machine check exception occurred |
0x1 |
A corrected machine check exception occurred |
0x2 |
A corrected platform error occurred |
0x3 |
A non-maskable interrupt (NMI) error occurred |
0x4 |
An uncorrectable PCI Express error occurred |
0x5 |
A generic hardware error occurred |
0x6 |
An initialisation error occurred |
0x7 |
A BOOT error occurred |
0xD |
SCI-based GHESv2, an ACPI generic hardware error source |
0xE |
Baseboard management controller error information |
0x10 |
Device driver error source |
0x11, 0x12 |
Arm synchronous external abort, Arm SError interrupt |
Debugger commands for a WHEA dump
| Command | What it gives you |
|---|---|
!analyze -v |
The usual analysis, and often a decoded summary of the error |
!errrec <parameter 2> |
The WHEA_ERROR_RECORD structure – the specific component and address |
!whea |
Additional WHEA state |
!errpkt |
The WHEA error packet |
The codes that arrive with this one
| Code | Published meaning | Where it sends you |
|---|---|---|
0x00000124 |
A fatal hardware error, reported through WHEA | The error record, then the component it names |
0x00000122 |
An internal error in WHEA itself | A vendor PSHED plug-in, or the firmware’s error record or injection implementation |
0x0000009C |
A fatal machine check exception | Replaced by 0x124 from Windows Vista; still issued in two narrow circumstances |
0x00000080 |
A hardware malfunction has occurred | No parameters; remove recently added hardware or drivers, check memory modules match |
0x000000B9 |
Nothing is published beyond the name | !analyze on the dump is the only documented step |
0x00000122 points at software, not silicon
This one is worth separating out. WHEA_INTERNAL_ERROR is an error inside the error-handling machinery, and Microsoft attributes it to a bug in a vendor-supplied PSHED plug-in, or to the firmware’s implementation of error records or error injection. So a 0x122 is a reason to update firmware and platform software, not to start pulling memory. Getting this the wrong way round wastes a lot of parts.
0x0000009C is not simply the old version of 0x124
It was replaced by 0x124 from Windows Vista onwards, but it is still issued on current Windows in two documented circumstances: when WHEA is not fully initialised, and when all the processors that rendezvous have no errors in their registers. Both mean the machine could not produce a usable error record, which is itself information – it usually points at the very earliest part of start-up, or at firmware.
When the evidence runs out
- Read the WHEA-Logger entries in the System log across the whole period, not just the crash. Corrected errors accumulate before an uncorrectable one arrives.
- Test with the minimum viable hardware: one memory module, no add-in cards, integrated graphics if the platform has it.
- Swap the power supply if you have a known-good one. It is not in Microsoft’s cause list, but an unstable rail produces symptoms that read as processor and memory faults.
- Try a clean current build of Windows on a spare disk. If it crashes there too, software is ruled out cheaply.
- Where parameter 1 is 0x10, treat it as a driver problem and work the driver rather than the hardware.
Microsoft’s advice for 0x00000080 is short and worth following literally: remove any hardware or drivers that were recently installed, and make sure all memory modules are of the same type. There are no parameters to read on that code, and it says so.
Every code this article covers
| Code | What it points at | Source |
|---|---|---|
0x00000124 |
WHEA_UNCORRECTABLE_ERROR: a fatal hardware error occurred, using error data provided by the Windows Hardware Error Architecture. Parameter 1 is the error source type and parameter 2 is the address of the WHEA_ERROR_RECORD, readable with !errrec | Microsoft Learn |
0x00000122 |
WHEA_INTERNAL_ERROR: an internal error in WHEA itself, attributed to a bug in a vendor-supplied PSHED plug-in or to the firmware implementation of error records or error injection | Microsoft Learn |
0x0000009C |
MACHINE_CHECK_EXCEPTION: a fatal machine check exception. Replaced by 0x00000124 from Windows Vista onwards, and on later Windows issued only when WHEA is not fully initialised or when all processors that rendezvous have no errors in their registers | Microsoft Learn |
0x00000080 |
NMI_HARDWARE_FAILURE: a hardware malfunction has occurred. No parameters, and Microsoft states the exact cause is difficult to determine; the published advice is to remove recently installed hardware or drivers and confirm all memory modules are the same type | Microsoft Learn |
0x000000B9 |
CHIPSET_DETECTED_ERROR: Microsoft publishes the symbolic name and the statement that this bug check appears very infrequently. No description, parameters or cause is given, so it is treated here as an unattributed platform error to investigate with the dump | not published by the vendor |
Confirm the fix worked
- Firmware is at stock settings with no memory profile enabled, and the machine has completed a sustained load test.
- The WHEA-Logger source in the System log records no new errors over several days.
- Windows Memory Diagnostics and the manufacturer’s own diagnostics both pass.
- System firmware and chipset drivers are at the versions currently published for that model.
- If a component was replaced, the decoded error record no longer names it.
Questions people ask about this
Can software cause 0x00000124?
Microsoft’s cause list ends with ‘less likely, but possible: a driver is causing the hardware to fail with this bug check’. So it is unlikely but not impossible, and parameter 1 of 0x10 is a device driver error source outright. Treat hardware as the strong prior, not a certainty.
Which component is failing?
Read the error record. Parameter 2 is its address and !errrec decodes it in WinDbg, naming the bank, address or device. Guessing from the stop code alone is how people replace three good parts before finding the bad one.
I undervolted for thermals. Is that a cause?
Microsoft names over-clocking, heat, cooling failure and defective memory or processors. Undervolting is not on the list, but it is a departure from stock that can destabilise a machine, so remove it while you test alongside everything else.
Does 0x0000009C mean my Windows is out of date?
No. It is still issued on current Windows in two narrow cases – WHEA not fully initialised, and no errors present in the registers of the processors that rendezvous. Both mean no usable error record could be produced.
Should I reinstall Windows?
Not for this code. Trying a clean build on a spare disk is a reasonable one-hour experiment to rule software out, but replacing the installation on the existing disk is not a fix for a hardware error report.
