Fix it now
DFS Replication hit an error talking to a named partner over RPC. The entry carries the partner name, its address and an error code, and that code is the diagnosis: reach the partner on TCP 135 and the dynamic range, then confirm authenticated RPC works between the two computer accounts.
Test-NetConnection <partner-fqdn> -Port 135
nslookup <partner-fqdn>
dcdiag /test:DFSREvent
dcdiag /test:NetLogons
w32tm /query /status
- Open the DFS Replication log and read the error code and string inside the 5002 entry. Microsoft’s worked example of the neighbouring 5008 carries Error 1722, the RPC server is unavailable, which is what a blocked path looks like from here.
- If TCP 135 answers but replication never starts, the RPC dynamic range is the suspect. Microsoft’s port list gives DFS Replication as TCP 135 plus a random high port between 49152 and 65535, with a further entry on TCP 57222.
- If the error is an access or authentication failure, compare the two clocks and confirm both computer accounts are healthy before touching the network.
- Once the path is open, make both members re-read their configuration with
dfsrdiag PollAD, then watch for Event ID 5004 against that partner.
Restarting the DFS Replication service re-establishes sessions and re-reads configuration. It does not open a firewall, fix DNS or correct a clock, so do it last rather than first.
If 5004 appears and the backlog drains, you are done. If not, the next section separates the four events that look identical in a log and point in different directions.
Why it happens
DFS Replication is a remote procedure call application. Each member contacts the endpoint mapper on TCP 135 on its partner, is handed a dynamically allocated port, and holds a session open over it. Microsoft’s own service and port requirements list DFS Replication as TCP 135 plus a randomly allocated high port in the 49152 to 65535 range, with an additional entry on TCP 57222. Authentication is between the two computer accounts, so a working connection needs name resolution, an open path, clocks close enough for Kerberos, and healthy machine accounts.
Event 5002 is the general failure of that conversation. Microsoft publishes it as EVENT_DFSR_CONNECTION_ERROR, “The DFS Replication service encountered an error communicating with the partner member”, and the entry carries the replication group, the partner’s full and short names, its IP address, and an implementation-specific error code and string. That last pair is the part worth reading. Everything else in the entry tells you where; only the status tells you why.
The neighbours divide the space. 5008 is published as the partner being unreachable or the DFS Replication service not running there, and Microsoft’s own example of it carries Error 1722. 5012 is the partner not recognising the connection or the replication group configuration, which is a directory problem rather than a network one. 6016 is different in kind and is the one most often misread: it is DFS Replication failing to write its own configuration back into Active Directory, not a partner event at all.
The RPC dynamic port range is blocked between subnets
You have this one if TCP 135 answers, replication never starts, and members on the partner’s own subnet work normally.
- Permit TCP 135 plus the documented dynamic range, 49152 to 65535, from each member to its partners.
- Check whether the environment already relies on the fixed DFS Replication port in Microsoft’s list, TCP 57222, and permit that too if so.
- Test from the failing member’s subnet. A path that works from the server proves nothing.
- Restart DFS Replication on both ends and watch for Event ID 5004.
DFS Replication can be pinned to a single static RPC port with dfsrdiag StaticRPC. Microsoft’s port list carries the DFS Replication entry on TCP 57222, but the dfsrdiag syntax is not on a current documentation page, so read it back with dfsrdiag /? before relying on it.
Name resolution points at an address the partner no longer holds
You have this one if nslookup on the partner’s fully qualified name returns nothing, or an address the server has not used since it was rebuilt.
- Remove stale host records and let the partner re-register.
- Clear the resolver cache on the failing member with
ipconfig /flushdns. - Run
dcdiag /test:DNSon both machines and act on what it reports. Microsoft notes this test is not run by default and must be requested explicitly. - Re-test, then poll with
dfsrdiag PollAD.
Authentication between the two computer accounts is failing
You have this one if The status inside the 5002 entry is an access or authentication failure rather than a transport error.
- Compare the clocks with
w32tm /query /statuson both machines. Microsoft’s own security check treats a skew of 300 seconds or more as fatal to Kerberos. - Run
dcdiag /test:CheckSecurityError /ReplSource:<partner>, which checks the skew, the permissions on every naming context, and connectivity to the SYSVOL and NETLOGON shares in one pass. - Confirm both computer objects exist and are enabled in the directory.
- Restart DFS Replication and re-test.
The configuration has not replicated, so the partner does not recognise the connection
You have this one if Event 5012 rather than 5002, shortly after a connection or replication group was created or changed.
- Force the change out from the domain controller where it was made, then let it converge.
- Make each member re-read it with
dfsrdiag PollAD. - Recheck the DFS Replication log for 5004 on both ends.
DFS Replication cannot write its own configuration to the directory
You have this one if Event 6016 rather than any of the connection events. Nothing is wrong with the partner.
- Read the entry as what Microsoft publishes it as: the service failed to update the configuration in Active Directory and will retry periodically.
- Check that the member can reach a writable domain controller and that its computer account is healthy.
- Run
dcdiag /test:MachineAccounton the server logging it.
Chasing a 6016 through the firewall is the classic wasted afternoon on this error. It is a directory write, not a partner conversation.
Full reference
What each event actually says
| Event | Microsoft’s published meaning | Where it sends you |
|---|---|---|
| 5002 | The service encountered an error communicating with the partner member; the entry carries the partner, its address and an error code and string | Read the status. It is the diagnosis |
| 5008 | Failed to communicate with the partner for the replication group; can occur if the host is unreachable or the service is not running there | The partner itself, or the path to it |
| 5012 | Failed to communicate with the replication partner; the partner did not recognise the connection or the replication group configuration | Directory replication, then re-poll |
| 5014 | Not published by Microsoft. Documentation names it alongside 5004 in a stop-and-resume cycle | Treat it as the connection dropping, with 5004 as the recovery |
| 6016 | The service failed to update the configuration in Active Directory and will retry periodically | The directory, not the partner |
Ports, from Microsoft’s own list
| Port | Protocol | What it is |
|---|---|---|
| 135 | TCP | RPC endpoint mapper |
| 49152 to 65535 | TCP | The randomly allocated high ports RPC uses on Windows Server 2008 and later |
| 57222 | TCP | The DFS Replication entry in Microsoft’s port table |
| 445 | TCP | SMB, which the rest of a domain controller’s work needs |
| 88, 389, 464, 3268, 3269, 53 | TCP and UDP | The remainder of the Active Directory set, needed on a DC for everything other than DFS Replication |
Tests worth running before you change anything
| Command | What Microsoft says it does |
|---|---|
dcdiag /test:DFSREvent |
Validates DFS Replication service health by checking DFSR event log warning and error entries from the past 24 hours |
dcdiag /test:NetLogons |
Validates that the user running it can connect to and read the SYSVOL and NETLOGON shares without security errors |
dcdiag /test:SysVolCheck |
Reads the Netlogon SysVolReady registry value. It does not check whether the shares are accessible |
dcdiag /test:CheckSecurityError /ReplSource:<dc> |
Checks time skew, naming context permissions, SYSVOL and NETLOGON connectivity, and the Access this computer from the network privilege |
dcdiag /test:MachineAccount |
Checks the computer object exists, sits in the right container, and carries the right flags and service principal names |
dfsrdiag PollAD |
Makes the member re-read its configuration from the directory |
SysVolCheck and NetLogons are easy to confuse and they answer different questions. Microsoft states plainly that SysVolCheck reads a registry value and does not test share access, so a passing SysVolCheck on a DC whose SYSVOL share is unreadable is not a contradiction.
When the path is open and it still fails
- Confirm the partner is a member of the replication group you think it is. A connection removed on one side and left on the other produces a steady 5012 that no firewall change will clear.
- Check whether the service on the partner is running at all. Microsoft’s published meaning for 5008 names that case explicitly.
- On a newly promoted domain controller, look for 4612 and 4614 instead. Microsoft documents a promotion in which SYSVOL never finished initial replication because the seeding partner named in the registry could not be reached, and the DC then refuses to advertise.
- Look at what else on the box holds RPC. A security product that intercepts or proxies RPC between servers produces this event with no firewall rule to find.
- If several members fail at once and nobody touched the network, look at the directory rather than the wire.
The seeding partner on a new domain controller
Microsoft documents a specific failure on a freshly promoted DC where SYSVOL replication never starts because the parent computer recorded for the seeding operation is unreachable. The value lives under the DFSR parameters key, in the SysVols branch, under Seeding SysVols for the domain, as a Parent Computer value naming the source DC. The documented fixes are to repair name resolution or connectivity to that machine, or to point the value at an available source domain controller. If your 5002 or 5008 is on a DC that has never once served SYSVOL, check there before anything else.
Reading a backlog rather than guessing
A replication group that is failing and a replication group that is merely slow look the same from the event log. The DFS Replication diagnostic tool reports the number of files waiting to move between a named sending and receiving member, and a backlog that falls after you make a change is the only evidence that matters. Take a reading before you touch anything, so that the second reading means something.
When a licence is the actual fix
Nothing here is a licensing fault, and the usual answer is a firewall rule, a DNS record or a configuration poll, which costs nothing. A purchase becomes relevant in one case: the partner named in the event is a domain controller you have already decided to retire, or one whose Windows Server release you are consolidating away from, and nursing the connection along is postponing the replacement rather than fixing anything. That replacement needs a licensed Windows Server. Microsoft publishes the edition difference plainly – Windows Server 2025 Standard permits two virtual machines plus one Hyper-V host per licence, Datacenter permits unlimited virtual machines plus one Hyper-V host per licence, and the number of users either edition serves is governed by your client access licences. Arco can work through the count with you and check what your existing agreement already covers before you spend anything.
Every code this article covers
| Code | What it points at | Source |
|---|---|---|
Event ID 5002 |
The DFS Replication service encountered an error communicating with the partner member. The entry carries the replication group, the partner’s names and address, and an error code and string | Microsoft Learn |
Event ID 5008 |
The service failed to communicate with the partner for the replication group. Microsoft names two causes: the host is unreachable, or the DFS Replication service is not running on it | Microsoft Learn |
Event ID 5014 |
Logged when communication with a partner stops after an error; Microsoft’s documentation names it beside 5004 in a stop-and-resume cycle but does not publish the message text | not published by the vendor |
Event ID 5012 |
The service failed to communicate with the replication partner because the partner did not recognise the connection or the replication group configuration | Microsoft Learn |
Event ID 6016 |
The service failed to update the configuration in Active Directory and will retry periodically. This is a directory write failure, not a partner communication failure | Microsoft Learn |
Confirm the fix worked
- Event ID 5004 appears in the DFS Replication log for the partner that was failing.
Test-NetConnection <partner> -Port 135returns true from the failing member’s own subnet.dcdiag /test:DFSREventreports no DFS Replication errors or warnings in the last 24 hours.dcdiag /test:NetLogonssucceeds, which proves the SYSVOL and NETLOGON shares are readable rather than merely flagged ready.- A policy change made on one domain controller appears on the other.
Questions people ask about this
Which ports does DFS Replication need?
TCP 135 for the endpoint mapper plus a randomly allocated high port between 49152 and 65535, and Microsoft’s port table also lists a DFS Replication entry on TCP 57222. On a domain controller you additionally need the rest of the Active Directory port set for everything else the server does.
Is 5002 a replication conflict?
No. A conflict is a content decision taken after two members have compared what they hold. Event 5002 is the conversation between them failing, so nothing has been compared yet.
Why does 6016 appear alongside the connection events?
Because a member that cannot reach a domain controller properly will fail both to talk to its partner and to write its own configuration back. They share a cause but not a remedy: 6016 is fixed in the directory, and no firewall rule clears it.
Will restarting the service fix it?
It re-establishes sessions and re-reads configuration, so it clears transient cases and connections the service had already given up on. It does not open a port, correct a record or fix a clock, so treat it as the last step.
Does any of this cost money?
No. Every tool involved ships with Windows Server or the remote administration tools, and the fixes are configuration on your own network.
