Skip to content

Writing and rehearsing an AD forest recovery plan

Build an Active Directory forest recovery runbook from Microsoft's guide: clean room, first DC restore, SYSVOL, FSMO, RID pool, krbtgt resets and yearly drills.

Florian Amette7 min read

Forest recovery is the procedure nobody wants to run, which is why most organisations have never run it. It is the last resort when Active Directory itself cannot be trusted: DCs encrypted by ransomware, a forest-wide destructive change, or a compromise so deep that no DC can be declared clean. Microsoft's Active Directory Forest Recovery Guide describes the procedure in detail. Reading it during an incident, with the business down and executives on the bridge, is the wrong time to find out that your DSRM passwords are in a vault that authenticates against AD.

This guide turns Microsoft's procedure into a runbook you own, adapted to your forest, and a rehearsal programme that proves it works. It assumes the backups from protecting AD backups from ransomware exist, and it goes deeper than the overview in the assessment, backup and recovery pillar.

Measure: what the plan must know

A recovery plan is mostly inventory. Collect and keep offline, on paper and on encrypted media outside AD:

  • Forest topology: every domain, its DCs, their sites, IP addresses and OS versions, and which DCs hold FSMO roles and the global catalog.
  • Backup map: for each domain, which DCs are backed up, in what format, where the immutable copy lives, and how to reach it without AD.
  • Credentials: DSRM passwords for backed-up DCs, break-glass accounts for hypervisors, storage and backup consoles that do not depend on AD, and the offline vault procedure.
  • Dependencies: DNS, DHCP, AD CS, Entra Connect, AD FS, PAM, and the order in which applications must come back.
  • People: named owners for each phase, with deputies, and out-of-band communication that does not rely on Exchange or Teams signed in through the compromised AD.
PowerShell
# Snapshot of topology to print and store with the plan
Get-ADForest | Select-Object Name, ForestMode, SchemaMaster, DomainNamingMaster, Domains, GlobalCatalogs
Get-ADDomain | Select-Object DNSRoot, DomainMode, PDCEmulator, RIDMaster, InfrastructureMaster
Get-ADDomainController -Filter * | Select-Object HostName, Site, IPv4Address, OperatingSystem, IsGlobalCatalog, OperationMasterRoles
(Get-ADObject "CN=Directory Service,CN=Windows NT,CN=Services,$((Get-ADRootDSE).configurationNamingContext)" -Properties tombstoneLifetime).tombstoneLifetime

An empty tombstoneLifetime means the old 60-day default for forests created before Windows Server 2003 SP1. Forests created later default to 180 days.

Decide: is this a forest recovery?

Put the decision criteria in the plan so the incident commander is not improvising:

SituationResponse
Deleted objects, healthy DCsAD Recycle Bin or authoritative restore of specific objects
One or a few DCs failed, others healthyMetadata cleanup and re-promote new DCs
All DCs in one domain lost, rest of forest healthyRecover that domain following the forest recovery procedure for it
DCs encrypted or wiped, or persistent Tier 0 compromise that cannot be scopedFull forest recovery into a clean environment

Also decide which backup is clean. With a long attacker dwell time, the newest backup may contain persistence: rogue admins, modified AdminSDHolder, SID History, malicious GPOs, golden ticket-capable keys. The plan should name who makes that call and what evidence they use, typically forensic timelines from your event forwarding pipeline.

The recovery procedure

The order below follows Microsoft's guide. Keep your runbook aligned with the current version of that guide, not with this summary.

1. Build the clean room

Build an isolated network with no route to production, clean hypervisor or hardware, clean installation media verified by hash, and admin workstations built from trusted images. Everything that follows happens here. The restored forest joins production only when you decide it is clean.

2. Restore the first writable DC in the forest root domain

Restore one writable DC, preferably one that was a global catalog and DNS server, from the chosen backup. If the original DC's server or VM can be reused, boot it into DSRM and restore system state as below. If it is gone, do a bare-metal recovery from a full server backup onto isolated hardware first, then continue from DSRM.

PowerShell
# On the DC to be restored, reboot into DSRM
bcdedit /set safeboot dsrepair
shutdown /r /t 0

# After logging on with the DSRM account
wbadmin get versions -backupTarget:E:
wbadmin start systemstaterecovery -version:09/20/2026-02:00 -backupTarget:E: -authsysvol -quiet

This is a non-authoritative restore of AD DS with an authoritative restore of SYSVOL. The -authsysvol switch marks this DC's SYSVOL as the primary copy. If the restore method does not support that switch, set msDFSR-Options to 1 on CN=SYSVOL Subscription,CN=Domain System Volume,CN=DFSR-LocalSettings,CN=<DC>,OU=Domain Controllers,<domain DN> before restarting.

Before leaving DSRM, stop the DC from waiting for replication partners that no longer exist:

PowerShell
Set-ItemProperty -Path 'HKLM:\SYSTEM\CurrentControlSet\Services\NTDS\Parameters' `
  -Name 'Repl Perform Initial Synchronizations' -Value 0 -Type DWord
bcdedit /deletevalue safeboot

3. Take control of the domain

Once the DC is up in normal mode, with DNS pointing at itself:

  • Seize every FSMO role held by DCs that will not be recovered:
PowerShell
Move-ADDirectoryServerOperationMasterRole -Identity 'DC01' `
  -OperationMasterRole SchemaMaster, DomainNamingMaster, PDCEmulator, RIDMaster, InfrastructureMaster -Force
  • Metadata cleanup for every other DC in the domain. Deleting the DC's computer object in Active Directory Users and Computers, or its server object in Active Directory Sites and Services, performs the cleanup on current Windows Server versions, with ntdsutil metadata cleanup as the fallback. Also remove their DNS records and any NS records and delegations pointing at them.
  • Raise the RID pool by 100,000 on the RID Manager object (rIDAvailablePool on CN=RID Manager$,CN=System), and invalidate the current RID pool on the restored DC. Both steps prevent duplicate SIDs for objects created after the backup was taken.
  • Reset the restored DC's computer account password, and reset each side of every trust (netdom trust ... /resetOneSide) once the partner domain is recovered.

4. Reset credentials

  • Reset the krbtgt password twice. With a single DC there is no replication to wait for, but let the first reset complete before the second. This invalidates every ticket signed with the old keys. The mechanics are in krbtgt password rotation.
  • Reset the built-in Administrator and every Tier 0 account password, disabling any account you cannot vouch for.
  • Plan the reset of all user passwords, service account passwords and gMSA keys if the backup may have been stolen. The backup contains every hash as of backup time.

5. Recover the other domains, then rebuild

Repeat steps 2 to 4 for one writable DC in each domain, working top-down from the forest root. Domains can be recovered in parallel once the root is up. Follow the guide's global catalog steps for your topology, because users cannot log on until a GC is available. After that, promote fresh DCs from clean media. Do not restore more backups. Reconnect sites and replication gradually, watching for reinfection.

6. Clean up before reconnecting

Remove persistence found by the forensic team, re-run a PingCastle assessment and an ACL review, and verify privileged group membership before any production traffic reaches the recovered forest.

Verify

Every recovery checkpoint needs an objective test:

PowerShell
dcdiag /v /c /e /f:C:\Recovery\dcdiag.txt
repadmin /replsummary
repadmin /showrepl * /csv > C:\Recovery\showrepl.csv
Get-ADDomainController -Filter * | Select-Object HostName, OperationMasterRoles, IsGlobalCatalog

# SYSVOL: 4602 = authoritative SYSVOL initialized on the first DC; 4604 = non-authoritative sync done on new DCs
Get-WinEvent -FilterHashtable @{ LogName = 'DFS Replication'; Id = 4602, 4604, 4614 } -MaxEvents 10
Get-SmbShare | Where-Object Name -in 'SYSVOL','NETLOGON'

A DC with 4614 but no 4604 is still waiting for initial SYSVOL replication and will not share SYSVOL. That usually means the first DC's SYSVOL was never marked authoritative.

Rehearse

  • Tabletop, twice a year. Walk through the runbook with every named owner. Check that each credential and phone number is where the plan says it is, and that nobody's step silently depends on AD, email or SSO.
  • Technical drill, yearly. Restore real backups from the immutable copy into an isolated lab and execute the runbook to the point of a working forest with fresh DCs. Time every phase.
  • Update after every drill. Each drill produces fixes: missing drivers, wrong DSRM passwords, undocumented FSMO placement, scripts that assumed a specific OS build. Commercial forest recovery tools (Semperis ADFR, Quest Recovery Manager for AD) can automate much of this, but they need the same rehearsal.

What it breaks

  • Data loss to the backup point. Every change after the chosen backup is gone: new users, password changes, group changes, computer joins. Machines joined or re-keyed since then may lose their secure channel and need Test-ComputerSecureChannel -Repair or a rejoin.
  • Kerberos and sessions. Resetting krbtgt twice invalidates all tickets. Every user and service re-authenticates, and long-running services may need restarts.
  • Hybrid identity. Entra Connect sees the restored directory as a large change set. Stage the sync server, review the pending exports, and watch the accidental deletion threshold.
  • Drill risk. A restored DC accidentally connected to production reintroduces old passwords and lingering objects. Physically or logically isolate the lab, and check it before every drill.
  • Time. Realistic forest recovery for a multi-domain forest takes days, not hours. The drill timings, not wishful RTOs, belong in your business continuity plan.

Related reading: protecting AD backups from ransomware makes sure a clean backup exists, defining Tier 0 lists everything that must come back with the forest, and the assessment, backup and recovery area collects the full series.

Frequently asked questions

When is a full forest recovery needed rather than restoring a single domain controller?

When no DC in the forest can be trusted or kept running: ransomware has encrypted or wiped DCs, an attacker has held Domain or Enterprise Admin long enough that persistence cannot be ruled out, a bad schema change has replicated everywhere, or every DC in a domain is gone. If healthy DCs remain and the problem is a deleted object or a single failed DC, use the Recycle Bin, an authoritative restore of specific objects, or a normal re-promotion instead.

Why does Microsoft recommend a non-authoritative restore for the first DC in each domain?

In a forest recovery, every other DC is removed and rebuilt, so there is no replication partner whose changes need to be overwritten. A non-authoritative restore of AD DS is enough and avoids needlessly bumping version numbers on every object. SYSVOL is the exception: it must be restored authoritatively on that first DC, by using the authsysvol option or setting msDFSR-Options to 1, so DFSR treats it as the primary copy and does not wait for partners that no longer exist.

How often should we rehearse forest recovery?

Run a tabletop walkthrough of the runbook at least twice a year and after major changes such as a new domain, a DC operating system upgrade or a new backup product. Run a technical drill that restores real backups into an isolated lab at least once a year. Record how long each phase took; the timings are what you give leadership as a realistic recovery time objective, and they show which steps need automation.

Writing and rehearsing an AD forest recovery plan

Related guides

Assessment, Backup & Recovery

Protecting Active Directory backups from ransomware

Design AD backups that survive ransomware: system state per domain, DSRM passwords, immutable and offline copies, a backup system outside the AD it protects.

Intermediate
Assessment, Backup & Recovery

AD assessment, backup and forest recovery

Run recurring AD posture assessments against CIS and Microsoft baselines, protect Tier 0 backups, and rehearse forest recovery before you need it for real.

Foundation