AZ-400: CI/CD Pipeline Architecture

A delivery pipeline is a control system for turning source changes into production behavior. Good pipelines do more than run builds. They preserve artifact identity, apply tests and security checks, control who can promote changes, separate environments, limit blast radius, and collect enough evidence to decide whether a release should proceed or roll back.

The current AZ-400 places the largest share of its current blueprint on designing and implementing build and release pipelines. The broader Azure path show where the DevOps role fits, but AZ-400 pipeline questions reward architecture and operational judgment rather than memorizing one YAML syntax.

A strong design starts by asking what must be trusted at each stage: source, build environment, dependencies, artifact, deployment identity, target environment, approval evidence, and production telemetry.

Source control is the pipeline’s first trust boundary

Branches, pull requests, required reviews, protected paths, and commit history define how code becomes eligible for automation. If anyone can bypass review or rewrite history on critical branches, later pipeline controls inherit untrusted input. Repository permissions should therefore align with the sensitivity of the systems being changed.

Branch strategy should support the team’s delivery model rather than become ceremony. Trunk-based development, release branches, or other patterns can work when the pipeline makes version, promotion, and rollback behavior clear. The important point is traceability from a deployed artifact back to an approved source state.

Build environments should be repeatable and isolated

Build agents execute code from the repository and often have access to package feeds, signing keys, artifact stores, and deployment systems. They are high-value targets. Ephemeral or well-controlled agents, minimal permissions, trusted base images, patched toolchains, and network restrictions reduce the chance that a build compromise reaches production.

Repeatability matters for diagnosis. If two builds from the same source can silently use different dependencies or tools, troubleshooting and provenance become harder. Pinning versions, caching carefully, and producing build metadata help teams understand exactly what created an artifact.

Tests and gates should correspond to real risk

Unit, integration, security, compliance, performance, and acceptance tests serve different purposes. Running every possible test on every small change may slow feedback unnecessarily, while skipping critical checks creates dangerous shortcuts. Pipeline architecture should place fast feedback early and more expensive validation where it meaningfully reduces release risk.

The broader CI/CD provides the general concept. AZ-400 candidates should go further by connecting gates to environment promotion, approvals, deployment strategy, and operational evidence.

Artifacts should be immutable and traceable

Once an artifact passes build and validation, later stages should promote the same artifact rather than rebuild it from source. Rebuilding can introduce dependency or environment changes that invalidate earlier testing. Artifact repositories, versioning, checksums, signatures, and provenance data support consistent promotion.

Traceability should connect artifact version, source commit, build run, dependencies, test results, approvals, and deployment target. That evidence supports incident investigation, rollback, audit, and change review.

Environment boundaries control blast radius

Development, test, staging, and production environments exist to separate risk, but they only help when identities, secrets, data, and deployment permissions are isolated appropriately. A pipeline identity that can change every environment from one credential creates a large blast radius.

Environment-specific configuration should be managed deliberately. Secrets should come from protected stores rather than source control, and configuration changes should be versioned or otherwise auditable. Teams also need realistic test environments; a staging system that differs fundamentally from production may provide false confidence.

Deployment strategy changes how failures surface

Rolling, blue-green, canary, feature-flag, and other strategies distribute change risk differently. A canary release limits initial exposure but requires telemetry good enough to compare new and old behavior. Blue-green deployment can make rollback fast but may require additional capacity and careful data compatibility.

Choosing a strategy depends on application architecture, user impact, database changes, traffic controls, and recovery requirements. AZ-400 questions often reward the strategy that provides the needed safety and feedback rather than the one with the most automation.

Approvals should add judgment, not delay for its own sake

Manual approval can be valuable when a release has business, regulatory, or operational risk that automated tests cannot fully evaluate. It becomes waste when approvers lack evidence or merely click through every release. Effective approvals present change scope, test results, risk indicators, and deployment plan.

Automated policy gates can enforce repeatable requirements such as vulnerability thresholds, required artifacts, infrastructure-policy compliance, or successful tests. The pipeline should distinguish what machines can verify consistently from what still needs human judgment.

Rollback needs both technical and data compatibility

Deploying the previous application binary does not always restore the previous state. Database migrations, message formats, external integrations, and irreversible side effects can make rollback complex. Pipeline design should address backward compatibility, migration strategy, backup or recovery points, and the point at which roll-forward is safer than rollback.

Current AZ-400 design emphasizes this systems view: delivery safety depends on application design, data, infrastructure, identity, and observability working together.

Production telemetry closes the delivery loop

A pipeline should not declare success only because deployment commands returned successfully. Health checks, error rate, latency, business transactions, infrastructure signals, and user feedback show whether the release achieved the expected outcome. Automated rollback or halt conditions are safest when they rely on signals that actually represent service health.

Post-deployment evidence also improves future delivery. If a release passed every pre-production test but failed in a predictable production pattern, teams should decide whether a new test, environment change, or monitoring gate can catch that class of issue earlier.

AZ-400 pipeline architecture is therefore a chain of trust and evidence: approved source creates a repeatable artifact, tests and policies validate it, environments and identities limit blast radius, deployment strategy controls exposure, and telemetry verifies the result. Automation is valuable because it makes that chain repeatable—not because every step must be automatic.

Pipeline identity deserves its own threat model. Build and release systems often have permission to read source, fetch packages, sign artifacts, access secrets, modify infrastructure, and deploy to production. Those credentials should be separated by stage and environment where practical, rotated, audited, and protected from untrusted pull-request code or user-controlled scripts.

Dependency management is part of pipeline architecture as well. External packages, container bases, actions, extensions, and build tools can introduce supply-chain risk. Version pinning, trusted feeds, integrity checks, software composition analysis, and controlled update processes make builds more reproducible and reduce the chance that a third-party change silently enters production.

Infrastructure as code should travel through the same evidence model as application code. Plans or previews can show intended changes before deployment, policy checks can reject unsafe configurations, and drift detection can reveal manual changes outside the pipeline. Treating infrastructure changes as code makes rollback and peer review more consistent, but only if state and secrets are handled safely.

Database and schema delivery often requires special sequencing. A backwards-compatible migration may need to deploy before new application code, remain valid while old instances drain, and be cleaned up later. Pipelines should model those dependencies explicitly rather than assuming application rollback can reverse every data change.

Promotion policy should also account for emergency fixes. Hotfix paths can be faster while still preserving review, artifact identity, and production evidence. An emergency process that bypasses source control or builds directly on a production server may restore service once but create an untraceable state that causes future failures.

Finally, pipeline metrics should measure flow and quality together. Lead time, deployment frequency, failure rate, rollback rate, queue time, test duration, and recovery time reveal different bottlenecks. Optimizing only for speed can increase instability; optimizing only for control can make delivery too slow to be useful. Good AZ-400 architecture balances both.

Reusable pipeline templates can improve consistency when they encode approved build, test, security, and deployment patterns. The risk is hidden coupling: a template change can affect many repositories at once. Versioned templates, change review, testing, and clear ownership help teams gain reuse without turning a shared pipeline library into an uncontrolled single point of failure.

Artifact retention should support rollback and audit needs without keeping every build forever. Teams need policies for which releases, logs, test results, signatures, and provenance records are retained and for how long. Retention should match incident-response, compliance, and recovery expectations.

Pipeline troubleshooting should follow the same evidence chain as design. First identify which stage failed, then determine whether the cause is source, dependency, agent, test, artifact, credential, environment, deployment target, or post-release health. Changing multiple stages at once makes diagnosis harder and can hide the original failure.

Feature flags can separate deployment from release. Code may be deployed safely while user exposure is controlled independently, allowing teams to test infrastructure and application health before enabling a capability broadly. Flags still require ownership and cleanup; long-lived flags create hidden branches in production behavior and complicate testing.

Release orchestration should also respect dependencies between services. In distributed systems, one service may need to remain backward compatible while another is upgraded, and contract changes may require versioned APIs or staged rollouts. A pipeline that deploys components in the wrong order can create failures even when each component passed its own tests.

Pipeline design should include observability for the pipeline itself. Queue delays, agent failures, flaky tests, artifact-store errors, approval bottlenecks, and deployment retries are operational signals. If the delivery system is unreliable, teams lose confidence in automation and begin bypassing it, weakening the control model AZ-400 expects candidates to understand.

  • img