Skip to content

AD assessment, backup and forest recovery

Run recurring AD posture assessments against CIS and Microsoft baselines, protect Tier 0 backups, and rehearse forest recovery before you need it for real.

Florian Amette7 min read

Hardening controls degrade silently. A Group Policy setting drifts, a service account picked up Domain Admin membership during an incident nobody rolled back, a new trust arrived with a migration project and was never revisited. Assessment, backup, and recovery rehearsal are the governance layer that catches drift and guarantees that when — not if — something goes wrong, you have both the evidence to know it happened and a tested path back to a clean forest.

This post covers running and interpreting posture assessments, comparing configuration against Microsoft's Security Compliance Toolkit and CIS Benchmarks, treating LAPS and tiering as recurring checks rather than one-time projects, protecting Tier 0 backups, and rehearsing full forest recovery.

Posture assessment tools

Three tools cover most of what you need, and they are complementary rather than redundant:

ToolFocusOutput
PingCastleAD-specific risk scoring across Stale Objects, Privileged Accounts, Trusts, AnomaliesRisk score 0–100 per category, trend-over-time reports, an actionable checklist with direct links to remediation guidance
Purple Knight (Semperis)AD, Entra ID, and Okta security indicators mapped to MITRE ATT&CKIndicator-level pass/fail with severity, useful for tracking specific technique coverage
Microsoft's own tooling (Microsoft Defender for Identity secure score, AD Security Assessment via Microsoft Services)Native integration with Defender/Entra telemetryRecommendations surfaced directly in the Defender portal

Run PingCastle (or Purple Knight) as a scheduled task rather than a one-off:

PowerShell
# PingCastle, healthcheck mode, scheduled via Task Scheduler
PingCastle.exe --healthcheck --server dc01.domain.com --level Full

Reading the score: don't chase 100/100 — some findings are false positives for your environment or accepted risk (documented, e.g., a legacy application requiring an older protocol). What matters is:

  1. Critical findings trend down release over release, never up.
  2. Any new Trusts or Privileged Accounts finding gets triaged within days, not the next quarterly cycle — these categories change fastest and are the most commonly attacker-relevant.
  3. Findings are tied to a ticket and an owner, not just re-reviewed and re-accepted indefinitely.

Baseline comparison: Microsoft SCT and CIS Benchmarks

Posture scanners tell you about AD-object-level risk (stale accounts, trust misconfiguration, delegation). They generally do not fully replace a Group Policy baseline comparison, which tells you whether your actual GPOs match a vetted security baseline.

  • Microsoft Security Compliance Toolkit (SCT): ships baseline GPO backups for each supported Windows Server version, plus Policy Analyzer, a GUI/CLI tool that diffs your production GPOs against the baseline and flags every setting that differs.
  • CIS Benchmarks for Windows Server / Active Directory: an independently maintained, more prescriptive baseline; many regulated environments require CIS compliance specifically rather than (or in addition to) Microsoft's baseline.

Workflow:

PowerShell
# Export current GPOs for comparison
Get-GPO -All | ForEach-Object { Backup-GPO -Guid $_.Id -Path "C:\GPOBackups" }

Then load both the exported backups and the SCT/CIS baseline GPO backups into Policy Analyzer, and generate a diff report. Track exceptions (settings you deliberately deviate from) in a document with a justification and an owner — auditors and future you will both want it.

Re-run this comparison whenever Microsoft publishes a new SCT baseline (typically aligned with each Windows Server feature update) and at least twice a year regardless.

LAPS and tiering as recurring checks, not one-time projects

Both LAPS (Local Administrator Password Solution / Windows LAPS) and administrative tiering (Tier 0 and privileged access) rot if treated as a deployment project rather than an ongoing control:

PowerShell
# Verify LAPS is actually rotating passwords, not just installed
# Windows LAPS (msLAPS-* attributes)
Get-ADComputer -Filter * -Properties msLAPS-PasswordExpirationTime |
  Where-Object { $_.'msLAPS-PasswordExpirationTime' -lt (Get-Date).AddDays(-35).ToFileTime() } |
  Select-Object Name
# Legacy Microsoft LAPS, only where its schema extension (ms-Mcs-AdmPwd*) is installed
Get-ADComputer -Filter * -Properties ms-Mcs-AdmPwdExpirationTime |
  Where-Object { $_.'ms-Mcs-AdmPwdExpirationTime' -lt (Get-Date).AddDays(-35).ToFileTime() } |
  Select-Object Name

# Verify no unexpected accounts have landed in Tier 0 groups
Get-ADGroupMember "Domain Admins" -Recursive | Select-Object Name, SamAccountName
Get-ADGroupMember "Enterprise Admins" -Recursive | Select-Object Name, SamAccountName

Schedule both checks weekly at minimum. Group membership drift — a service account added to Domain Admins during an incident and never removed, a vendor granted temporary access that outlived the engagement — is one of the most common ways tiering silently fails, and it is invisible unless you actively look. Correlate against the 4728/4732/4756 events so you know when and by whom the drift happened, not just that it exists.

Protecting Tier 0 backups

A domain controller's system state backup is, functionally, a copy of every credential in the domain: the NTDS database with all password hashes, and the krbtgt key that signs every Kerberos ticket. Protect it accordingly:

  • Take system state backups, not just file-level or VM snapshots, so authoritative and non-authoritative restore both work correctly:
PowerShell
wbadmin start systemstatebackup -backupTarget:E: -quiet
  • Store backups offline, immutable, or air-gapped from production. If a Domain Admin-equivalent credential (or ransomware that obtained one) can reach and delete your backup store, it isn't a backup — it's a second copy of the same asset the attacker already owns. Immutable object storage (write-once, retention-locked) or true offline/air-gapped media are both acceptable; a backup share joined to the same domain, reachable by the same privileged accounts, is not.
  • Restrict backup operator rights — Backup Operators can read the entire NTDS database via a system state backup, making that group functionally equivalent to Domain Admin for confidentiality purposes. Audit its membership as strictly as Domain Admins.
  • Test restorability, not just backup completion. A backup job reporting success tells you it wrote data; it does not tell you the data restores.

Rehearsing forest recovery

Microsoft publishes a detailed AD Forest Recovery Guide covering the exact order of operations for recovering from a scenario where the forest cannot be trusted (ransomware, malicious schema change, or the loss of enough DCs to break replication/quorum). Do not read it for the first time during an actual incident.

High-level order of operations (see Microsoft's guide for the full procedure):

  1. Identify the last known-good, malware-free system state backup for one DC per domain, prioritizing a Global Catalog holder in the forest root domain.
  2. Isolate the environment — disconnect from the network/pause replication — before restoring, so a still-active adversary or corrupted replication partner cannot re-infect the restored DC.
  3. Restore the first DC in each domain (forest root first) in Directory Services Restore Mode (DSRM) as a non-authoritative restore of AD DS with an authoritative restore of SYSVOL (wbadmin start systemstaterecovery ... -authsysvol, or set msDFSR-Options to 1 on the DC's SYSVOL subscription for DFSR). Every other DC will be rebuilt, so there is no replication partner whose changes need to be overwritten.
  4. Reset the krbtgt password twice in each domain as part of recovery, respecting replication convergence between the two resets — this invalidates every Kerberos ticket issued before recovery, including any forged by an attacker.
  5. Clean up metadata for any DCs that will not be recovered, using ntdsutil, so stale DC objects don't cause replication or DNS problems:
Text
ntdsutil
metadata cleanup
connections
connect to server <survivingDC>
quit
select operation target
list domains
select domain <n>
list sites
select site <n>
list servers in site
select server <n>
quit
remove selected server
quit
quit
  1. Rebuild and re-promote remaining DCs from clean media once the first DC and metadata cleanup are verified good.
  2. Re-enable network connectivity and replication progressively, monitoring for reinfection signals before fully reconnecting the estate.

Rehearse this in an isolated lab at least annually. Recovery guides read differently than they execute — DSRM password issues, unexpected FSMO role placement, or unfamiliarity with ntdsutil syntax under pressure are exactly the kind of friction you want to discover during a drill, not during a ransomware incident at 2 a.m.

What this breaks

Nothing directly — this is governance and disaster-recovery work, not a production-facing control. The cost is time and process discipline: scheduled assessment runs, backup storage budget for immutable/offline retention, and periodic recovery drills that consume lab infrastructure and staff hours. None of it changes day-to-day authentication or access behavior for users.

Combine this with Kerberos hardening for the krbtgt rotation cadence outside of recovery scenarios, and Auditing, logging, and detection so that when posture assessment flags a finding, you also have the logs to determine whether it was ever exploited.

Frequently asked questions

How often should we run a PingCastle or Purple Knight assessment?

At minimum quarterly, and after any significant change such as a new trust, a forest/domain migration, or onboarding an acquired company's AD. Many teams run it monthly as a scheduled task and track the trend of the risk score over time — a single point-in-time score matters less than whether the score is improving or degrading.

Why does DC backup need to be offline or immutable rather than just backed up normally?

A domain controller's system state backup contains the NTDS database, including every password hash and the krbtgt key, in a form an attacker can extract offline. If backups sit on the same network as production, reachable by a Domain Admin-equivalent credential, an attacker who compromises the domain can also delete or tamper with the backups, denying you clean recovery. Offline, immutable, or air-gapped backup storage ensures a ransomware or destructive attack that reaches your DCs cannot also destroy your last clean copy.

Do we need to reset the krbtgt password during a normal patching cycle, or only during recovery?

Routine krbtgt rotation (twice, respecting the default 10-hour ticket lifetime gap) is a recurring hygiene task independent of recovery — see Kerberos hardening for that cadence. During an actual forest recovery from backup or after a suspected compromise, krbtgt must be reset twice as part of the recovery procedure itself, because every Kerberos ticket issued before the compromise was discovered must be invalidated.

AD assessment, backup and forest recovery

Related guides

Assessment, Backup & Recovery

Protecting Active Directory backups from ransomware

Design AD backups that survive ransomware: system state per domain, DSRM passwords, immutable and offline copies, a backup system outside the AD it protects.

Intermediate
Assessment, Backup & Recovery

Writing and rehearsing an AD forest recovery plan

Build an Active Directory forest recovery runbook from Microsoft's guide: clean room, first DC restore, SYSVOL, FSMO, RID pool, krbtgt resets and yearly drills.

Advanced
Assessment, Backup & Recovery

PingCastle assessment: from AD report to action plan

Run a PingCastle health check on Active Directory, read the four risk scores correctly, triage findings into owners and sprints, and track progress over time.

Foundation