Exchange DR requires immutable backups and geo-redundancy

Exchange DR requires immutable backups and geo-redundancy

Exchange DR requires immutable backups and geo-redundancy

Exchange disaster recovery needs two separate protections: backups that cannot be changed or deleted, and copies stored in another location. One protects the backup from tampering. The other protects it from a site or region failure. A recovery plan that uses only one of them has a clear gap.

I start with this because Exchange high availability can create false comfort. A Database Availability Group, or DAG, keeps copies of mailbox databases on several Exchange servers. Microsoft describes a DAG as a system for database-level recovery when a server, network, or database copy fails. It can support a fast failover, but it is still part of the live Exchange system.

That difference matters. If an attacker gains control of the Exchange environment, damage may spread to database copies. If a bad change reaches several copies, the DAG may preserve the wrong state very well. A DAG helps with uptime. It does not replace an isolated backup.

Two protections with different jobs

An immutable backup is a backup that cannot be changed or erased during its protection period. The control may be supplied by the storage system or backup platform. The exact setting varies, but the result should be the same: an attacker with access to normal backup operations cannot alter the protected copy.

This protects against ransomware, mistaken deletion, and silent changes to backup files. It also helps preserve an older recovery point. That point may be needed when a problem is found long after it began.

Geo-redundancy means keeping another recovery copy in a different site or region. The second location must not depend on the same failed power system, storage system, network path, or building. A second disk shelf in the same room is a copy. It is not useful geo-redundancy.

These controls solve different problems. Immutability protects the trustworthiness of a backup. Geographic separation protects access to a backup when the main site is lost. One cannot stand in for the other.

Microsoft’s Exchange guidance supports Exchange-aware, VSS-based backups. VSS is the Windows service that coordinates snapshots with applications. Exchange-aware backup software uses the Exchange VSS writer so database files and transaction logs are handled as a working set. A simple file copy is not the same thing as a supported Exchange backup.

Where the recovery path starts

A sound plan begins with a known recovery point. That means knowing which backup contains the needed mailbox data and which logs belong with it. It also means knowing where the backup can be restored if the primary site is unavailable.

The recovery target may be a full mailbox database, a replacement server, or a recovery database. A recovery database is a special Exchange database used to mount restored data without replacing the live mailbox database. Microsoft documents using it to extract mailboxes or items with the New-MailboxRestoreRequest cmdlet.

That path is useful because it separates data recovery from production service. A restored database can be checked in an isolated location. Selected mailbox data can then be moved into current mailboxes. This does not remove every risk. Database damage, missing logs, version limits, permissions, and storage problems can still stop or limit recovery.

The plan also needs a clear order. A typical recovery design answers these questions:

  • Which copy is the protected source?
  • Who can release or unlock it?
  • Where is the alternate recovery site?
  • Which Exchange version can mount the restored data?
  • How are mailbox contents restored to users?
  • How is the result checked before normal service resumes?

These are operational questions, not paperwork. Under pressure, unclear ownership can delay a recovery even when the backup itself is sound.

The limit of replication

Replication is often described as if it were backup. It is not. A DAG can reduce downtime when an Exchange server or database copy fails. It can also support site resilience when copies are placed across sites and the design is configured correctly.

Still, replication usually follows the active system. A deleted mailbox, damaged database state, or unwanted change may reach another copy. Replication can preserve availability while failing to preserve a clean historical point.

That is why the backup layer needs its own controls. The protected copy should be separate from the Exchange servers and from routine administrator access. The remote copy should be separated from the primary site. Restore access should be tested with the same care used for backup jobs.

There is also a limit to the word “immutable.” A storage lock does not prove that the backup can be restored. It does not prove that all required logs are present. It does not prove that the restore process fits the Exchange version or the available hardware.

A backup is useful only when it can produce usable data. Restore tests show whether the whole chain works. They can expose expired credentials, missing encryption keys, weak network links, wrong paths, and unclear recovery steps.

A practical recovery view

I think about Exchange DR as three layers.

The first layer keeps service available. DAG copies and failover procedures belong here.

The second layer preserves clean recovery points. Exchange-aware backups, immutable storage, and retention controls belong here.

The third layer makes recovery possible after a site loss. Geo-redundant storage, alternate infrastructure, documented access, and restore tests belong here.

A weakness in one layer does not disappear because another layer is strong. Fast failover cannot replace a clean backup. An immutable local backup cannot replace a remote copy. A remote copy cannot help much if nobody can restore it.

The exact design depends on the Exchange version, database size, recovery time target, and available infrastructure. Microsoft documents the technical recovery tools, but no document can predict every damaged database or failed storage path. Some recoveries may restore the database but lose recent mail. Others may recover only selected mailboxes or items.

That uncertainty belongs in the plan. A recovery method should state what it can restore, what it may lose, and what must be tested. It should never promise that every corrupt database will open cleanly.

The main decision is simple. Exchange DR needs a live-service plan and a trusted-data plan. The live-service plan may use DAG copies and site failover. The trusted-data plan needs immutable backups stored with geographic separation.

That design gives Exchange two independent recovery paths. One handles system failure. The other handles damage that replication may repeat.

Exchange Admin Notes carries the same practical focus: Exchange Server recovery tips, migration notes, and administration shortcuts for IT professionals. That focus keeps disaster recovery tied to the steps that make restored data usable.