Fix it now
Event ID 55 is NTFS saying the file system structure on the disk is corrupt and unusable and asking you to run chkdsk. On SAN or iSCSI storage the media is almost never the problem: a write was acknowledged and never landed. Stop writing to the volume, scan it online, and find the storage event that caused it.
fsutil dirty query X:
chkdsk X: /scan
Repair-Volume -DriveLetter X -Scan
- Move or shut down anything writing to the volume first. If it holds a clustered role or virtual machines, get them off before you touch the file system.
- Look for Event ID 98 next to the 55. It names the volume and tells you plainly that it needs to be taken offline for a full chkdsk, which means an online scan is no longer enough.
- Where an offline pass is needed, schedule it:
chkdsk X: /ffor the repair, orchkdsk /f /rwhere you also want bad sectors located and readable information recovered. - Read the System log for the same window. Microsoft groups 55 and 98 with events 129, 153 and 157 for a reason: the storage event is usually the cause and the NTFS event the consequence.
- Take a backup or storage snapshot of the volume as it stands before any offline repair.
On a clustered disk or a Cluster Shared Volume, put the resource into maintenance mode first and run the repair from the coordinator node, not from wherever you happen to be logged on.
If the scan is clean, the dirty flag is clear and the storage events have stopped, you can stop here. The next section explains how a healthy disk produces a corrupt file system.
Why it happens
NTFS keeps a journal so that a crash leaves the file system consistent: metadata changes are written to the log before they are applied, and the log is replayed on mount. That guarantee rests on one assumption – when the storage says a write is durable, it is durable. Break the assumption and the journal becomes fiction.
There are ordinary ways to break it on shared storage. A path fails mid-write and the array acknowledges data it never committed. A thin-provisioned LUN runs out of real capacity while the host still believes it has space. A controller cache is enabled without the power protection it assumes and the server loses power. In each case the disk surface is perfect and the structure on top of it is not.
This is why the two events matter in a specific order. Event 55 says the structure is corrupt and unusable and asks for chkdsk. Event 98 goes further: it names the volume and its device path and states that it needs to be taken offline for a full chkdsk, telling you to run CHKDSK /F locally or REPAIR-VOLUME <drive:> locally or remotely from PowerShell. A 55 without a 98 may still be repairable online; a 98 is the point at which it is not.
Writes were lost when a storage path dropped
You have this one if NTFS events in the same hour as event 129, 153 or 157, or an iSCSI session loss.
- Correlate the NTFS events against the System log for the same window and identify the path failure.
- Fix the path first. Repairing the file system while the path is unreliable just produces new damage.
- Once storage is stable, run
chkdsk X: /scanand act on what it queues. - Restore from backup anything the repair could not reconcile rather than trusting a partially repaired structure.
The array ran out of real capacity behind a thin LUN
You have this one if Damage on several volumes at once, all from the same pool, with the host still reporting free space.
- Ask the storage team for the pool’s real utilisation, not the LUN’s presented size.
- Free or add capacity to the pool before touching the file system.
- Repair each affected volume in turn afterwards, and restore what cannot be repaired.
- Set alerting on pool utilisation so the next occurrence is a warning rather than a corruption event.
Caching promises the hardware cannot keep
You have this one if Damage appears after an unplanned power loss or a controller reset, on a volume where write-cache buffer flushing was turned off.
- In Device Manager, open the disk device and check the Policies tab.
- Leave write caching on only where the controller has battery or flash-backed cache, and treat turning off buffer flushing as something you do only where the vendor explicitly supports it.
- Confirm the array’s cache battery or supercapacitor is healthy.
- Confirm the server is on a UPS with a tested shutdown path.
The volume mounts and logs the same event at every boot
You have this one if Event 137 recurs on a volume that otherwise behaves, and no storage event accompanies it.
- Have the default transaction resource manager clean its metadata at the next mount:
fsutil resource setautoreset true X:\– the trailing backslash is part of the documented syntax. - Reboot so the reset happens at mount time.
- Confirm the event does not return over the next two boots.
- If it does, run an offline pass in a maintenance window.
Microsoft publishes no message text for Ntfs event 137, so treat this as a low-cost thing to try rather than a diagnosis.
Full reference
Repair commands and what each one costs
| Command | Effect |
|---|---|
chkdsk X: /scan |
Online scan of an NTFS volume. Does not dismount it |
chkdsk X: /spotfix |
Spot fixing on the volume – a brief offline pass for what the scan queued |
chkdsk X: /scan /forceofflinefix |
Bypasses online repair and queues every defect found for an offline fix. Must be used with /scan |
chkdsk X: /f |
Fixes errors on the disk. The volume must be locked, so it is unavailable for the duration |
chkdsk X: /f /r |
Repairs and, in addition, locates bad sectors and recovers readable information |
chkdsk X: /scan /perf |
Uses more system resources to scan as fast as possible, at the cost of everything else running. Must be used with /scan |
Repair-Volume -DriveLetter X -Scan |
The PowerShell equivalent named in event 98 itself, usable remotely |
fsutil dirty query X: |
Reports whether the volume is flagged as needing a repair |
fsutil resource setautoreset true X:\ |
Has the default transaction resource manager clean its transactional metadata at the next mount |
Which events Microsoft groups with these
| Event | What it says |
|---|---|
| 55 | The file system structure on the disk is corrupt and unusable; run chkdsk on the volume |
| 98 | The named volume needs to be taken offline for a full chkdsk; run CHKDSK /F locally or REPAIR-VOLUME from PowerShell |
| 129 | A reset was issued to the storage device |
| 153 | An I/O at the named block address was retried |
| 157 | The disk was surprise removed |
| 137 and 130 | Seen on NTFS volumes, with no message text published by Microsoft |
Microsoft’s own data corruption guidance lists 50, 55, 98 and 140 as the NTFS corruption indicators, and puts the storage events beside them. Reading the file system event alone, without the storage events from the same window, is how a path problem gets treated as a disk problem.
A full offline repair takes the volume out of service for as long as it takes – on a file server with millions of small files, hours – and it can discard data it cannot reconcile. Back up or snapshot the volume as it stands before you start. On a clustered disk or CSV, suspend the resource so the volume is in maintenance mode, and repair only from the coordinator node.
Repairing in the right order
- Stop the writes. Move roles and virtual machines off, or shut them down.
- Take a backup or storage snapshot of the volume in its damaged state.
- Run the online scan and read what it reports.
- Fix the storage-side cause, if the System log shows one.
- Run the offline pass only if the scan or an event 98 says it is needed, in a window sized for the file count rather than the volume size.
- Re-run the scan afterwards and confirm the dirty flag is clear.
- Restore anything the repair detached or discarded, from a backup taken before the damage.
When the volume is clean and the events keep coming
- Check whether the same LUN is presented to more than one host without cluster arbitration. Two servers writing to one unshared LUN produce exactly this pattern.
- Check the array’s own event log for the same window; it may record a controller failover the host only saw as latency.
- Confirm the volume is not being snapshotted by two products at once, which shows up as intermittent damage on an otherwise healthy path.
- If the volume is a CSV, look for events 5120 and 5142 in the FailoverClustering log. A CSV that lost its path mid-write is the storage story behind the NTFS one.
Every code this article covers
| Code | What it points at | Source |
|---|---|---|
Event ID 55 |
NTFS: the file system structure on the disk is corrupt and unusable, and the volume needs chkdsk | Microsoft Learn |
Event ID 137 |
Logged by NTFS against a volume. Microsoft publishes no message text for this ID; where it recurs at every boot, having the default transaction resource manager clean its metadata at the next mount is a documented low-risk step | not published by the vendor |
Event ID 98 |
NTFS: the named volume needs to be taken offline to perform a full chkdsk; run CHKDSK /F locally or REPAIR-VOLUME from PowerShell | Microsoft Learn |
Event ID 130 |
Logged by NTFS against a volume, with no message text published by Microsoft. Microsoft’s own corruption indicator list is 50, 55, 98 and 140, so weigh those first | not published by the vendor |
Confirm the fix worked
fsutil dirty query X:reports the volume is no longer flagged.chkdsk X: /scancompletes with no damage found.- A week passes with no new Ntfs events on that volume.
- No storage resets, path losses or pool capacity warnings appear in the same period.
- A test restore from backup of a file on that volume opens correctly, proving the backup taken before the repair is usable.
Questions people ask about this
Will chkdsk lose data?
It can. Its job is to make the structure consistent, and where a file’s metadata cannot be reconciled it will detach or discard it. That is why you back the volume up before repairing it rather than after.
How long does a full offline repair take?
It scales with the number of files rather than the size of the volume, so a file server with millions of small files can run for hours. Plan the window generously and do not interrupt it.
Does this cost anything to fix?
No. chkdsk, Repair-Volume and fsutil all ship with Windows Server and no licence changes any of it. The only money involved is whatever the underlying hardware or capacity problem turns out to need.
Should I reformat instead?
Only when you have a restore you trust and the repair has failed. A reformat guarantees a clean structure and guarantees you are restoring from backup, so it is a decision about your restore, not about the file system.
The volume is a CSV. Does anything change?
Yes. Put the disk into maintenance mode so the volume is not in use, and run the repair from the node that owns it. Repairing a CSV from a non-coordinator node is not the same operation.
