Microsoft AI-200: Event-Driven AI Application Design

Event-driven and message-based integration is an explicit AI-200 objective. The current AI-200 exam from Microsoft expects developers to queue back-end work with Azure Service Bus, handle queues/topics/subscriptions/dead-letter queues, implement Event Grid workflows with filters/custom events/retries, and build Azure Functions with triggers and bindings.

AI workloads are a natural fit for these patterns because ingestion, embedding, document processing, evaluation, model calls, tool execution, and post-processing can be slow, bursty, or failure-prone. The design challenge is to preserve state and intent while separating the producer from the work.

The sections below focus on the semantics that matter for AI-200: commands versus events, idempotency, retries, poison messages, state, scaling, observability, and recovery.

Decide whether you are sending work or announcing a fact

Service Bus messages and Event Grid events can both move information, but they serve different ownership models. In practice, that means commands, queues, topics, subscriptions, events, filters, multiple consumers, and downstream state. Use a work queue when one or more workers own a durable unit of processing; use an event when publishers announce that something happened and consumers decide how to react. A document-upload event might notify several consumers, while the actual embedding job could be a durable queued command.

Treating every notification as a work command creates accidental coupling, while treating required work as a best-effort event can lose ownership. Useful evidence includes message contract, consumer responsibility, delivery semantics, retry path, and final state location. AI-200 directly tests both Service Bus and Event Grid, so expect scenarios that require choosing between them. Make that distinction explicit in diagrams.

Make consumers idempotent before relying on retries

In event-driven design, it helps to recognize that at-least-once delivery means a message or event may be processed more than once. The working parts are message IDs, business keys, state checks, upserts, deduplication, transactional boundaries, and retry behavior. Design the operation so repeating the same delivery does not corrupt state or create duplicate external effects. If a downstream side effect cannot be repeated safely, record a durable operation state before invoking it.

AI pipelines often duplicate documents, embeddings, notifications, or tool actions when retries are added after the code is already written. Validate the result with duplicate-delivery tests, state transitions, idempotency records, and consistent final results. The blueprint names retries and dead-letter handling; idempotency is the design principle that makes those features safe. Then resume from that state rather than guessing whether the first attempt succeeded.

Use dead-letter queues as an operating workflow, not a trash bin

For event-driven systems, first separate the design goal from the implementation detail: a dead-letter queue is useful only when the team can explain why messages land there and how they return to a valid processing path. The implementation normally spans delivery count, validation failures, expired messages, poison payloads, dependency failures, operator review, repair, replay, and retention. Classify failures into transient, permanent, and data-quality categories and define a different recovery action for each. Be careful with sensitive AI input when storing failed payloads; evidence should be useful without expanding data exposure.

Messages accumulate silently when no alert, owner, diagnostic context, or replay process exists. The evidence that matters is dead-letter rate, reason codes, preserved payload metadata, incident tickets, repair actions, and replay results. AI-200 explicitly calls out Service Bus dead-letter handling. A repaired message should still pass normal authorization and validation.

Separate orchestration state from model output

event-driven systems need durable state that does not depend on one message or one model response. Operationally, the design touches job records, workflow status, correlation IDs, intermediate artifacts, retries, timeout, cancellation, and final result. Persist the state needed to resume or explain a workflow, then let messages drive transitions. The message should reference durable state rather than being the only copy of it when the process has business value.

If state lives only in message bodies or function memory, a retry or deployment can make the workflow impossible to reconstruct. A review should look for state history, event timestamps, correlation IDs, transition logs, and final outcome. This is especially important for long-running AI processes built from Functions, queues, data stores, and model services. That also makes monitoring easier.

Use Event Grid filtering to keep consumers focused

Treat the event flow as an operating design rather than a feature checklist: publishers should emit meaningful events while subscriptions select the events each consumer actually needs. The system includes event types, subjects, custom attributes, filters, subscriptions, retries, and event handlers. Filter as close as practical to the subscription so consumers do not implement the same routing logic repeatedly. Keep event schemas stable enough for consumers to evolve independently.

Sending every event to every consumer increases cost, noise, and the chance that unrelated handlers react incorrectly. Prove the design with subscription configuration, filtered-delivery counts, handler logs, retry metrics, and unexpected-event tests. The current guide explicitly names filters, custom events, and retries. If a consumer needs the current resource state, fetch it from the authoritative store rather than trusting a stale event snapshot.

Use Functions for compact event handlers and APIs

Azure Functions work well when the trigger and execution fit a stateless, managed-scaling boundary. From there, the engineer or manager has to coordinate HTTP triggers, Service Bus triggers, Event Grid triggers, bindings, app configuration, secrets, deployment, and logging. Keep handlers focused on one transition or integration responsibility and move durable state to external services. Use queues to break a long pipeline into restartable steps rather than increasing a timeout.

A function becomes hard to recover when it hides several unrelated steps, long processing, and implicit state in one invocation. The most persuasive evidence is invocation logs, dependency traces, trigger metrics, retry behavior, and deployment versions. Functions, triggers, bindings, configuration, and deployment are direct AI-200 objectives. Each step should be safe to retry.

Scale on the right signal

event-driven AI systems should scale from backlog and processing demand rather than only CPU. In practice, that means queue depth, oldest-message age, event rate, concurrency, KEDA, worker throughput, model quota, and downstream limits. Choose scaling signals that reflect unfinished work and downstream capacity. Put a maximum on concurrency when downstream services have finite quotas or fragile rate limits.

Adding workers can make latency worse when a database, model quota, or external API becomes the bottleneck. Useful evidence includes backlog trend, processing duration, dependency saturation, throttle rates, and scaling events. AI-200 explicitly covers KEDA in Container Apps, which makes event-driven scaling a likely scenario. Scale-out should reduce backlog without destabilizing dependencies.

Secure the event path end to end

For event-flow security, assume that the producer, broker, consumer, and downstream service each need identity, authorization, data protection, and least privilege. The working parts are managed identities, RBAC, queue/topic permissions, event subscription permissions, Key Vault, network boundaries, payload sensitivity, and logging. Grant each workload only the send, listen, or management permissions it actually needs. Do not assume a trusted broker makes the payload trusted; consumers should still validate schema and authorization context.

One shared connection string can give many components broader broker access than their role requires. Validate the result with role assignments, managed identity use, denied-access tests, secret inventory, and broker audit logs. Key Vault and managed identities and cloud identity and least privilege support the security side of this design. Keep administrative identities separate from runtime identities.

Observability should explain flow, delay, and failure

Event-flow observability improves when you separate the design goal from the implementation detail: a healthy event-driven system needs visibility into both individual requests and aggregate backlog behavior. The implementation normally spans correlation IDs, distributed traces, queue depth, message age, retries, dead-letter counts, function duration, dependency latency, and business completion. Instrument the path so an operator can tell whether delay is at the producer, broker, worker, or downstream dependency. Include business-level signals such as ‘documents fully indexed’ rather than infrastructure metrics alone.

Teams often see low CPU and assume the system is healthy while messages age in a queue or fail repeatedly. The evidence that matters is OpenTelemetry traces, KQL queries, broker metrics, function logs, DLQ trends, and completion SLAs. OpenTelemetry and KQL are explicit AI-200 monitoring objectives. That distinction reveals partial failures.

Design recovery and replay before production

event-driven systems are resilient when they can safely replay work and reconstruct state after a failure or bad deployment. Operationally, the design touches retained messages or source events, durable input data, versioned code, idempotent consumers, backfill tools, dead-letter replay, rollback, and reconciliation. Define what can be replayed, from where, and how duplicate side effects are prevented. Practice a deployment that corrupts one stage, then recover without reprocessing unrelated work.

Without replay design, teams may restore infrastructure but still be unable to recover missed or corrupted AI processing. A review should look for tested replay runbooks, reconciliation counts, rollback results, and repaired state. General CI/CD fundamentals principles help because code versioning and rollback are part of reliable event-driven recovery. Recovery should be observable and auditable.

Production event flows should also prove how they recover after a dependency is unavailable for longer than a normal retry window. Validate dead-letter handling, replay order, duplicate tolerance, poison-message isolation, and the point at which an operator must intervene. A healthy dashboard is not enough if a backlog can grow silently while downstream work stops. For AI-200, this turns Service Bus, Event Grid, and Functions from separate products into one recoverable processing system with explicit delivery and state guarantees.

  • img