Fix it now
Microsoft publishes Event ID 1069 as a cluster resource in a clustered service or application having failed. Underneath a database availability group that means the Windows failover cluster broke, not Exchange. Fix the cluster first; everything Exchange does with database copies sits on top of it.
Get-ClusterNode
Get-ClusterGroup
Get-ClusterResource
Get-ClusterQuorum
Get-ClusterLog -Destination C:\Temp -TimeSpan 60
- Read the resource name and type from the 1069 entry. Microsoft’s own resolution starts there: check and correct problems with the application or service behind that resource.
- Check the witness. Confirm the server is up, the share exists and every member can reach it by name.
- Read the sequence in Event Viewer under Applications and Services Logs, then Microsoft, then Windows, then FailoverClustering.
- If a resource cannot start within reasonable time limits, Microsoft names the Pending timeout property on the resource as the value to adjust, after diagnosing why it is slow.
- Do not reseed databases or restart Exchange services while the cluster is broken. You will be reseeding into a group that cannot decide who owns anything.
Get-ClusterLog writes the detailed cluster log for the period you ask for. Sixty minutes around the incident is usually enough and is far easier to read than a full log.
If the cluster forms, quorum is intact and databases have mounted on their intended servers, you are done. If not, the next section covers quorum, the witness and the resource host process.
Why it happens
A database availability group is a Windows failover cluster with Exchange’s own manager layered on top. Exchange does not run its databases as traditional cluster resources, but it depends completely on the cluster for membership, for quorum and for the witness that breaks ties. When the cluster cannot establish quorum, nothing can decide which copy of a database should be active, so the safe outcome is that none of them are.
Event 1069 is published, and its wording is worth having exactly: cluster resource, in clustered service or application, failed. Microsoft’s symbolic name for it is RCM_RESOURCE_FAILURE, the source is Microsoft-Windows-FailoverClustering, and the documented cause is that a clustered resource failed to come online. The documented resolution is three things: correct problems with the application or service behind the resource, correct problems with cables or cluster-related devices, and adjust the resource’s properties – Pending timeout in particular – so there is enough time for the associated service to start.
The other two entries in this family are not published. One describes the process that hosts cluster resources being terminated and restarted; the other a node declining to start the cluster service because it cannot confirm its copy of the cluster configuration is current. Read what your own entries say. The second of those is a deliberate refusal rather than a fault, and clearing it means choosing a node you are confident about, which is a decision with consequences rather than a command to run casually.
The witness is unreachable or lacks the rights it needs
You have this one if An even-numbered group that will not hold quorum, and witness arbitration failures in the cluster log.
- Confirm the witness server is running and the share is reachable by name from every member.
- If the witness is not itself an Exchange server, add the Exchange Trusted Subsystem universal security group to its local Administrators group. Microsoft requires this so Exchange can create the directory and share.
- Enable the Windows Firewall exception for File and Printer Sharing on the witness. It uses SMB port 445.
- Configure it through Exchange with
Set-DatabaseAvailabilityGroup -Identity DAG1 -WitnessServer <server> -WitnessDirectory <path>. The witness cannot be a DAG member, must be in the same forest, and each DAG needs its own witness directory.
Quorum was lost because too many members are down
You have this one if Enough members are offline that the survivors cannot form a majority, and the cluster service has stopped on the rest.
- Bring members back rather than forcing quorum, if there is any prospect of doing so quickly.
- Check the quorum model with
Get-ClusterQuorumand confirm it matches the number of members you actually run. Microsoft names configuration changes such as adding nodes, when too few are online to achieve quorum in the new configuration, as a documented cause of quorum loss. - Where members are split across sites, understand which side is meant to survive an interruption before it happens, not during one.
- Once the members are back, confirm the cluster forms on its own before restarting anything in Exchange.
A resource fails or hangs and takes the host process with it
You have this one if The resource host process is terminated and restarted repeatedly, always around the same resource.
- Identify the resource from the entries and check what it depends on: storage, network or an agent.
- Isolate a repeatedly hanging resource into its own monitor process so its failures stop affecting others while you investigate.
- Check storage and network latency during the hangs. A resource that stops responding is usually waiting for something else.
- Where the resource is slow rather than broken, review its Pending timeout, which Microsoft names as the property that must allow enough time for the associated service to start.
The cluster name object in Active Directory is broken
You have this one if The cluster service will not start, and directory errors accompany the cluster entries.
- Check the computer object the cluster uses has not been disabled, moved or deleted, and that it still holds the rights it was created with.
- Confirm replication has carried any recent change to the domain controllers the members use.
- Repair the object through Failover Cluster Manager, which offers a repair action for exactly this, rather than recreating the cluster.
Full reference
What is published about these entries
| Entry | Published description |
|---|---|
Event ID 1069 |
Cluster resource in clustered service or application failed. Source Microsoft-Windows-FailoverClustering, symbolic name RCM_RESOURCE_FAILURE |
| Resource host process terminated and restarted | None found. Read your own entry for the resource named |
| Node declining to start the cluster service | None found. Treat it as a refusal to start with a configuration it cannot confirm is current |
Event ID 1177 |
The Cluster service is shutting down because quorum was lost. Useful context when 1069 arrives with a group going offline |
Forcing a cluster to start on a node you have not verified can bring the group up with an out-of-date configuration, and in the worst case allows a partitioned group to make conflicting decisions about which database copies are active. Establish which node holds the current configuration before you force anything, and never force two nodes on either side of a network partition.
Microsoft’s own resolution for a failed resource
- Check and correct any problems with the application or service associated with the resource. If it cannot start within reasonable time limits, diagnose and resolve what is making it slow.
- Check and correct any problems with cables or cluster-related devices.
- Adjust the properties of the resource in the cluster configuration, especially the Pending timeout value, so it allows enough time for the associated application or service to start.
- Verify by bringing the clustered service or application online through Failover Cluster Management and watching for further events.
Witness requirements worth checking every time
- The witness server cannot be a member of the DAG, and must be in the same Active Directory forest.
- One server can witness several DAGs, but each DAG needs its own witness directory.
- A non-Exchange witness server needs the Exchange Trusted Subsystem universal security group in its local Administrators group, granted before the DAG is created.
- Windows Firewall on the witness needs the File and Printer Sharing exception; the witness uses SMB port 445.
- Set-DatabaseAvailabilityGroup can reconfigure the witness server and directory where the witness lost its storage or somebody changed the share permissions.
- Do not host the witness on a member of the group, or on the same host or storage as the members.
Gathering evidence without guessing
Get-ClusterLog -Destination <path> -TimeSpan <minutes> writes the cluster’s own detailed log for the window you specify, on every node. That is the record of what the cluster decided and when, and it is considerably more informative than the event log summary. Collect it while the incident is fresh; the log rolls, and a cluster that has been flapping for a day will have overwritten the part you wanted.
When a licence is the actual fix
Nothing about a witness share, a quorum model or a cluster name object needs buying, and the fix for almost every cluster fault is configuration. The honest purchase case is the platform underneath. A witness or a member running a Windows Server version that no longer receives updates is not a base you can keep a mail service on, and Arco supplies Windows Server 2025 Standard licences and can check your core counts and CAL position. Two things to settle before you spend anything. Confirm your Exchange build is supported on the Windows version you are moving to, and note that Exchange Server 2016 and 2019 went out of support on 14 October 2025 with Exchange Server Subscription Edition as the supported successor, so the Exchange half of the platform may be the more pressing conversation.
Every code this article covers
| Code | What it points at | Source |
|---|---|---|
Event ID 1069 |
Cluster resource in clustered service or application failed. Source Microsoft-Windows-FailoverClustering, symbolic name RCM_RESOURCE_FAILURE; the documented resolution covers the service behind the resource, cabling and devices, and the resource’s Pending timeout | Microsoft Learn |
Event ID 1146 |
No description is published by Microsoft. In practice the process hosting cluster resources being terminated and restarted; read your own entry for the resource involved and isolate it into its own monitor while you investigate | not published by the vendor |
Event ID 1561 |
No description is published by Microsoft. Treat it as a node refusing to start the cluster service because it cannot confirm it holds the current cluster configuration, and establish which node does before forcing anything | not published by the vendor |
Confirm the fix worked
Get-ClusterNodeshows every member up andGet-ClusterQuorumreports the intended configuration.- The witness is reachable and being used, rather than merely configured.
- Databases have mounted on their intended servers and every copy is healthy.
- The failover clustering log records no further resource failures over the following day.
Questions people ask about this
Do I need to buy anything to fix a cluster fault?
Almost never. Quorum, witness and resource problems are configuration, and the tools ship with Windows Server. The exception is a server whose operating system version is out of support, and that is a platform decision rather than a fix for this event.
What does Microsoft actually say Event ID 1069 means?
That a cluster resource in a clustered service or application failed, with the documented cause being a clustered resource that could not come online. The resolution names the service behind the resource, cabling and cluster devices, and the resource’s Pending timeout.
Can the witness be a file share on any server?
It can be any reachable Windows server in the same forest that is not a DAG member, with its own witness directory. If it is not an Exchange server, the Exchange Trusted Subsystem group must be in its local Administrators group, and it needs the File and Printer Sharing firewall exception because the witness uses SMB port 445.
Should I destroy and rebuild the cluster?
No, not as a troubleshooting step. Rebuilding a cluster underneath a DAG is disruptive and rarely necessary, and most of the faults that tempt people towards it are witness or directory problems that repair cleanly.
Why did the databases dismount when only one node failed?
If losing that node cost the group its majority, the cluster stops rather than risk two halves each deciding they are in charge. The fix is a quorum configuration that survives the failures you actually expect.
