Microsoft AZ-104: Azure Backup and Site Recovery Design

Backup and disaster recovery are related, but they solve different failure problems. That distinction is central to Microsoft AZ-104. Azure Backup creates protected recovery points so data or workloads can be restored. Azure Site Recovery continuously replicates supported workloads and orchestrates failover so services can run from another location when the primary environment is unavailable.

The current AZ-104 study guide explicitly includes Recovery Services vaults, Backup vaults, backup policies, backup and restore operations, Azure Site Recovery configuration, failover to a secondary region, and monitoring through reports and alerts. Candidates should therefore think in terms of recovery objectives and failure modes rather than treating the services as interchangeable checkboxes.

Design starts with the failure you need to recover from

A deleted file, a corrupted database, a compromised virtual machine, a failed availability zone, and a regional outage are different events. A useful continuity design identifies which events matter, how much data loss is acceptable, how quickly service must return, and what dependencies have to recover together.

Backup is usually the right foundation when the primary problem is recovering data to an earlier point. Site Recovery is appropriate when the problem is keeping a VM-based workload available through an outage by replicating it and failing it over. Many production systems need both because replication can copy corruption or unwanted change while backup preserves historical recovery points.

Within the broader Microsoft Azure certifications, AZ-104 focuses on the administrator’s operating responsibility: configuring protection, verifying jobs, testing recovery, responding to alerts, and understanding what happens during failover and restore.

Azure Backup is policy-driven protection, not a one-time copy

Azure Backup uses policies to define schedule, frequency, and retention. Protected data is captured as recovery points and stored in a vault appropriate to the workload. Recovery Services vaults support Azure VMs and several other protected workloads; Backup vaults support additional newer scenarios. The exam expects candidates to recognize the correct management object rather than assuming every workload uses the same vault type.

Policy design should reflect business requirements. A short retention window might be enough for frequent operational mistakes but insufficient for audit or long-term recovery. A schedule that looks frequent on paper can still miss the recovery point objective if jobs fail and no one monitors them.

Microsoft also emphasizes security boundaries around vaulted data. Protection is useful only if an attacker or operator error cannot easily remove the same backups needed for recovery. Administrators should be familiar with vault security settings, access control, and recovery safeguards as part of a complete design.

Restore design matters as much as backup configuration

A successful backup job proves that data was captured, not that the application can be recovered within the required time. AZ-104 candidates should think through the restore target, dependencies, permissions, network configuration, and validation steps.

Restoring a complete VM, a disk, a file, or an application-aware workload can have different operational consequences. Restoring to an alternate location may be safer for investigation or validation because it avoids overwriting the current state. Cross Region Restore, where supported and configured, can provide another recovery option using data available in a paired region.

Testing is the difference between theoretical and operational recoverability. A continuity plan should include scheduled restore exercises, a definition of success, and documentation of the time and manual steps required. If nobody has tested the recovery path, the stated RTO is an assumption.

Site Recovery is about replication, failover, and failback

Azure Site Recovery is a managed replication and failover service for supported virtual-machine scenarios. For Azure-to-Azure protection, the target can be another region or, in supported designs, another availability zone. For supported on-premises sources, the target is Azure.

Site Recovery continuously replicates workload changes and maintains recovery points that can be selected during failover. A planned failover is used when the outage is expected and coordination is possible; an unplanned failover is used when the primary environment is unavailable. After the primary environment is ready again, workloads can be reprotected and failed back.

The design must include more than VM disks. Network mappings, target resource configuration, identity dependencies, DNS behavior, application connectivity, and operational access determine whether the recovered VM is actually useful.

Recovery plans coordinate applications rather than isolated machines

Multi-tier applications rarely recover correctly when every VM starts at the same time. Site Recovery recovery plans group protected machines and define the order in which recovery groups fail over. Machines within a group can run in parallel; groups can be sequenced so dependencies become available before the tiers that consume them.

A database tier, middleware tier, and web tier illustrate the point. The database should be available before middleware starts; middleware should be ready before users reach the front end. Recovery plans can also include automation and manual actions for steps that Site Recovery cannot infer from VM replication alone.

Test failovers are essential because they validate the orchestration without intentionally disrupting the primary workload. Microsoft recommends regular testing, and the result should be treated as evidence: did the application start, did dependencies resolve, did users reach it, and could the environment be cleaned up afterward?

RTO and RPO turn service features into architecture decisions

Recovery time objective is the acceptable time to restore service. Recovery point objective is the acceptable amount of data loss measured in time. Backup frequency, replication behavior, application consistency, recovery automation, and dependency complexity all influence whether those goals can be met.

A system with a four-hour RPO cannot be protected by a once-daily backup and still satisfy that requirement. A workload with a 15-minute RTO may need preplanned network mappings, automated recovery sequences, tested credentials, and rapid validation; simply having replicated disks is not enough.

Use these objectives to avoid overengineering as well. Not every development VM needs the same protection as a revenue-critical production service. Recovery design should match business impact, not a desire to apply the most expensive option everywhere.

Monitoring closes the loop between configuration and recoverability

Backup and replication systems produce jobs, alerts, health signals, and reports. An administrator needs to know when protection stops meeting the design, not discover the problem during an outage.

Microsoft is moving at-scale continuity management toward Azure Business Continuity Center, while older Backup center material remains part of the product history. Candidates should follow the current Microsoft Learn experience rather than memorizing a portal label that can change. The underlying responsibility remains: monitor protected items, failures, job state, compliance, and recovery readiness.

Alert response should be operationalized. A failed job needs an owner, a severity model, an escalation path, and a rule for when repeated failure means the workload is outside its recovery objective. Reports are useful only when they lead to action.

Vault configuration is part of the recovery design, not housekeeping. Administrators should understand storage redundancy choices, security settings, and features such as Cross Region Restore where they apply. A vault’s settings can change the recovery options available during a regional problem, so those decisions should be made before an incident rather than during one.

Backup and Site Recovery can also share a Recovery Services vault in some scenarios while protecting data in fundamentally different ways. Site Recovery stores replication configuration in the vault but the replicated workload data follows the service’s replication architecture; Azure Backup stores protected recovery data and points for restore. Keeping those roles separate prevents a common misconception that “the vault” itself means the same recovery mechanism is being used.

Security events add another design dimension. If ransomware or a compromised administrator damages the production environment, rapid replication may carry unwanted state to the recovery location. Historical, protected backup points provide a different recovery path. The continuity architecture should therefore consider cyber recovery as well as infrastructure failure.

Network recovery also deserves explicit practice. A recovered VM can be healthy while the application remains unreachable because subnets, network security groups, private endpoints, load balancers, DNS, or client routes still point to the primary environment. When reviewing a recovery design, trace a real user request all the way to the recovered service and verify every dependency on that path.

Prepare for AZ-104 by designing complete recovery stories

For each practice workload, write a recovery story from failure to restored service. What failed? What is the RPO and RTO? Which service protects it? Which vault and policy are involved? Where will the workload recover? What network and identity dependencies exist? How will you test the result?

Then introduce trade-offs. Would a backup restore be sufficient, or does the workload need Site Recovery? What happens if replication is healthy but DNS still points to the primary site? What if the recovery plan starts application tiers in the wrong order? What if backups exist but the retention policy no longer satisfies a compliance requirement?

AZ-104 does not require you to design every enterprise BCDR program from scratch. It does require administrator-level judgment about Azure Backup and Site Recovery. Candidates who can explain the failure model, protection method, recovery sequence, and monitoring evidence are much better prepared than candidates who only know how to click “Enable backup.”

  • img