Microsoft AZ-305: Designing Around Recovery Targets
When business leaders ask for a resilient Azure design, they often begin with a phrase such as “no downtime.” An architect has to turn that ambition into decisions that can be tested: which failures the system must tolerate, how much data may be lost, what restoration time is acceptable, and how much protection the organization can afford. Microsoft AZ-305 assesses that judgment across infrastructure, data, identity, governance and business continuity.
The exam is Designing Microsoft Azure Infrastructure Solutions, part of the Azure Solutions Architect Expert credential requirements. Microsoft’s April 17, 2026 outline assigns significant weight to infrastructure design, with additional domains for identity and monitoring, storage, and business continuity. The AZ-305 study resources page is the assigned ExamSnap reference; Microsoft’s study guide is the current authority for what is measured.
Consider a booking service that cannot afford to lose more than five minutes of confirmed transactions and should resume essential operations within an hour after a major failure. The five-minute limit is a recovery point objective (RPO); the hour is a recovery time objective (RTO). Those statements immediately rule out some simplistic approaches. A nightly backup would not by itself protect the most recent five minutes, and manual reconstruction of servers may make the one-hour deadline unrealistic.
Recovery architecture depends on the workload. Database replication, availability zones, regional replicas, backup policies and failover automation each address different risks. Replication can copy corrupt or mistaken data; backups may allow point-in-time restoration but take time to recover. The architect should combine mechanisms only when each one has a clear responsibility. The historical contrast in Microsoft AZ-303's retired architect path explains why legacy architecture notes should not be used as a current certification checklist. A specialized example of those recovery decisions is found in Microsoft AZ-120 SAP on Azure resilience, where SAP availability targets constrain the architecture.
An application distributed across zones may still fail if its only database, identity dependency or integration service sits behind a single fragile component. Trace the user request from DNS through load balancing, compute, data access and external dependencies. Identify where health probing, traffic steering and state management can fail. A load balancer cannot make an unhealthy backend healthy; a replica cannot help when the application has no safe failover procedure.
For a regional disaster, document how traffic will be rerouted, how data consistency will be evaluated, and what operators will verify before reopening the service. Rehearsals should include a credible rollback or failback plan. If the organization has never measured failover under production-like load, the recovery-time promise is still a hypothesis.
AZ-305 asks architects to choose relational, semi-structured and unstructured storage according to access patterns, scale, protection and cost. Transactional order data, analytical events and image attachments have different requirements. A relational service can preserve structured transactions; object storage can handle durable files; event streams and integration services may connect them without unnecessary coupling.
In a design exercise, begin with data invariants. Can users tolerate eventual consistency? Must an update be atomic across records? Is query latency predictable, or does demand spike unpredictably? What retention and deletion rules apply? These answers should come before a comparison of service pricing tiers. A technically available storage product can be a poor fit if its consistency or access pattern conflicts with the business workflow.
Access design includes Microsoft Entra identity, role boundaries, secrets management and authorization to both Azure and on-premises resources. Governance adds management groups, subscriptions, tagging, policy and compliance responsibilities. These are structural decisions. Adding security after selecting every application component tends to create exceptions, duplicated secrets and unmanaged access paths.
Imagine an organization operating workloads for several subsidiaries. A management hierarchy and policy baseline can provide common security constraints while preserving useful delegated control. The architect should decide which resources share subscriptions, which principals receive scoped roles, and how logs will reach the teams that must respond to incidents. The objective is a platform that remains manageable as the organization grows.
Azure virtual machines, containers, serverless functions and platform services serve different operational needs. Compare management overhead, portability, scaling behavior, latency and application constraints. An existing tightly coupled application might initially require VMs, while new asynchronous functions could suit event-driven components. A migration recommendation should account for skills, dependencies and the cost of changing application behavior, not only target hosting fees.
Messaging and API integration matter because systems rarely operate alone. A payment service, inventory platform and notification system can become fragile if every component waits synchronously for every other one. The architect may choose asynchronous messaging for resilient decoupling and an API gateway for controlled client access. Explain the resulting delivery guarantees, retry behavior and observability needs rather than stopping at product selection.
Monitoring design is not something an operations team should improvise after launch. Decide what events, metrics and traces demonstrate customer impact, which logs must be retained, and how alerts will travel to responsible staff. Include budgets and service limits: a design that meets performance targets but cannot be operated within its cost envelope is incomplete.
A strong AZ-305 practice case gives you imperfect requirements and asks you to defend a choice. Write down the assumptions, compare at least two viable architectures, identify what would make your recommendation change, and specify how the finished solution will be tested. The exam rewards that measured reasoning. The best answer is rarely the architecture with the largest number of Azure services; it is the one whose compromises are deliberate and visible.
