AZ-305: Business Continuity and Disaster Recovery
Business continuity design starts with a business statement, not a backup product. A workload has an acceptable amount of data loss, an acceptable recovery time, dependencies that must recover in a workable order, and failure scenarios the organization is willing or unwilling to tolerate. The current AZ-305 blueprint gives business continuity its own 15–20% domain because Azure architects must turn those requirements into recovery and high-availability designs.
The most important distinction is between keeping a service available and recovering it after disruption. High availability reduces interruption from expected component failures. Disaster recovery restores or relocates service after a larger failure. Backup protects recoverable copies of data. These capabilities overlap, but none automatically substitutes for the others.
The architect’s work is therefore to decide what must survive, what can be rebuilt, what data must be recoverable, which dependencies need coordinated recovery, and how the design will be tested. That judgment is central to Azure certifications at the architect level rather than a list of services to memorize.
Recovery Time Objective describes how quickly a service needs to be restored after disruption. Recovery Point Objective describes how much data loss is acceptable, usually expressed as a time window. These objectives force the architect to make trade-offs visible.
A workload with a very low RTO may require continuously available capacity, rapid automated failover, or pre-provisioned resources in another failure domain. A low RPO may require frequent replication or near-continuous data protection. Both usually increase cost and operational complexity.
The design should not default to the smallest possible RTO and RPO. If the business can tolerate four hours of downtime and one hour of data loss, an architecture built for seconds may waste money and create unnecessary operational burden. AZ-305 scenarios often reward matching the design to stated requirements rather than selecting the most resilient option in isolation.
Many disruptions should be absorbed without invoking a disaster-recovery plan. Compute instances can fail. A rack or fault domain can have problems. A zone can become unavailable. Application components can crash. High-availability design distributes workload and removes single points of failure so that routine faults do not create full service outages.
For compute, that can mean multiple instances, appropriate distribution across failure domains, load balancing, health probes, and the ability to replace unhealthy capacity. For relational data, it can mean replicas and a supported failover design. For unstructured data, redundancy choices affect both durability and availability.
The architect must understand the failure scope each mechanism protects against. Redundancy inside one region is not the same as regional disaster recovery. Multiple instances behind one dependency can still share a single failure point elsewhere in the stack.
Replication keeps another copy of data relatively current and can support rapid failover. Backup creates recoverable versions or copies that can protect against deletion, corruption, ransomware, logic errors, or the need to restore older state. Replicating bad data quickly is still bad data; replication alone is not a complete backup strategy.
Backups also need independent protection. If the same administrative mistake or compromised identity can delete both production data and its backups, the organization has not created a strong recovery boundary. Retention, immutability options where appropriate, access separation, and restore testing all contribute to backup quality.
AZ-305 candidates should be able to recommend backup and recovery approaches separately for compute, databases, and unstructured data. The details differ, but the reasoning is consistent: identify the data, the required recovery point, the restoration method, and the dependencies needed to make the restored workload usable.
A multi-region design is not useful if the application tier can fail over but identity, secrets, DNS, data, messaging, or network dependencies remain unavailable in the original region. Recovery planning must map the dependency graph rather than protect only the most visible servers.
Some dependencies can be active in multiple regions. Others may require replication and promotion. Some are global services with their own behavior. The architect should document which components are active-active, active-passive, recreated from infrastructure code, restored from backup, or manually reconfigured during recovery.
Ordering matters. A front end restored before its database, key store, or message system is ready may be technically online but functionally broken. Runbooks should express dependency order and verification points rather than simply listing resources.
Disaster-recovery diagrams often show the arrow from primary to secondary and stop there. Real operations eventually face a second problem: what happens after the original site or region is healthy again?
Failback can require data reconciliation, replication reversal, DNS or routing changes, capacity validation, and another controlled transition. An architecture that can fail over only once is incomplete. The design should define whether the secondary location can become the new primary indefinitely, whether failback is required, and what evidence proves that it is safe.
This is also a testability issue. If failback has never been rehearsed, the organization may discover hidden assumptions only after a real incident. Continuity is a lifecycle, not a one-way emergency switch.
A recovery plan that has never been exercised is a hypothesis. Testing reveals missing permissions, stale documentation, capacity assumptions, dependency gaps, DNS delays, application startup order, backup corruption, and human coordination problems that architecture diagrams do not show.
Tests should be designed around the requirement. A backup restore test proves that data can be recovered. A component failover test proves a narrower availability mechanism. A regional recovery exercise validates a much larger chain. Not every test needs to be destructive, but the organization should know which parts have been simulated and which have been proven under realistic conditions.
The generic business continuity governance perspective is useful here: ownership, testing cadence, and improvement matter alongside the Azure technical design.
Data consistency complicates recovery further. Some applications can tolerate asynchronous replication and a small RPO. Others require tightly coordinated writes across several components. The architect should know whether the recovery design can produce a transactionally usable state, not merely whether each individual data service has a copy in another location.
Recovery also needs measurable exit criteria. “The region failed over” is not enough if the application cannot authenticate users, process a transaction, read recent data, or reach a downstream dependency. Runbooks should include functional validation steps that prove the service is usable at the recovery site, along with clear authority for declaring the incident stabilized or initiating failback.
Resilience consumes resources. A fully active second region may provide excellent recovery time but can nearly duplicate some costs. A warm standby uses less capacity but needs scale-up and validation. A pilot-light design reduces steady-state cost further but increases recovery steps. Backup-and-rebuild can be economical for lower-priority workloads while producing a longer recovery time.
The architect should therefore connect cost to business priority. Critical transaction systems may justify expensive redundancy. Internal reporting or development services may tolerate slower restoration. Applying the same recovery tier to every workload wastes money and can make the overall program harder to operate.
Cost also includes people and complexity. A design with many custom failover scripts may look economical on infrastructure spend but create operational risk if only one engineer understands it.
Infrastructure as code can shorten recovery time, but only when the code, parameters, dependencies, and deployment credentials are available outside the failed environment. Rebuilding compute from a template does not restore application state by itself, and a template that depends on unavailable artifacts or secrets can fail at the moment it is needed most. Recovery design should therefore protect the automation supply chain as deliberately as the workload.
Capacity at the recovery location is another practical constraint. A secondary region may be technically supported but unable to provide the required SKU, quota, or service capacity during a widespread event. Architects should verify quotas, understand capacity assumptions, and decide whether critical resources must be pre-provisioned rather than requested only after disaster is declared.
Emergency conditions are not a reason to weaken security. The recovery environment needs access controls, secrets, certificates, logging, and network restrictions appropriate to production. Break-glass procedures may be necessary, but they should be controlled and auditable.
Identity itself is a dependency. Administrators need a way to access the recovery environment when normal systems are impaired. Applications may need managed identities, service principals, keys, or certificates that remain valid in the secondary location. If those dependencies are overlooked, technically restored infrastructure can remain unusable.
Recovery logs are also important. During a crisis, teams need evidence of what changed, who initiated failover, what completed successfully, and where the process diverged from the plan.
The strongest AZ-305 preparation starts with scenarios: one VM fails, a zone becomes unavailable, a database is corrupted, a region is lost, credentials are compromised, or an operator deletes critical data. For each scenario, state the required RTO and RPO, then identify which architecture mechanism addresses it.
This approach prevents service-name memorization from replacing design judgment. Azure offers many continuity capabilities, but the right recommendation depends on workload behavior, recovery objectives, data consistency, geographic requirements, cost, and operational maturity. The broad AZ-305 architecture perspective is useful precisely because resilience is a trade-off across these dimensions.
A candidate who can explain why a workload needs high availability, backup, replication, regional failover, or some combination of them is better prepared than one who simply recognizes product names. Business continuity is the proof that architecture can survive the failures the business has decided matter.
