Microsoft DP-420: Cosmos DB as an AI Memory Store

An AI assistant remembers that a customer asked for a delayed delivery, but a different customer later receives that preference in a suggested response. The underlying failure may have little to do with the language model. It may be a data-partitioning, identity or retrieval design error. Building reliable AI applications on Azure Cosmos DB means treating conversation memory as operational data with ownership, retention and cost constraints, not as an undifferentiated pile of JSON.

Microsoft revised DP-420 on October 6, 2026, toward Building AI Applications with Azure Cosmos DB and the Azure Cosmos DB AI Developer Associate role. Partitioning and efficient querying still matter, but retrieval-augmented generation and agent memory now belong in the preparation picture.

Start with the ownership boundary

Imagine a support application serving thousands of customer organizations. Every conversation entry has a tenant, a user, a session and a timestamp. The natural urge is to partition everything by session ID, because the assistant often loads the latest conversation. That may make a single-session read efficient while scattering tenant-level reporting and complicating retention operations. Another choice, tenant ID alone, risks concentrating activity for one enormous tenant. The right key follows measured access patterns and expected growth.

Sketch the actual reads before choosing a container: latest messages for a session, user-specific preferences, supporting document chunks, and operational audit events. Decide which belong together, which need different lifetimes and which require transactional consistency. A document database gives flexibility in structure, but flexibility does not remove the need for explicit tenancy rules or schema-version handling. Test whether queries touch one logical partition or fan out across many; the distinction affects both latency and request-unit consumption.

Retrieval has a security context

Semantic similarity is not permission. A vector search can surface a highly relevant chunk that the requesting employee should never see. Grounding pipelines must apply access constraints before sensitive material reaches a model, and the application must retain those constraints when documents are reindexed or user memberships change. Separate retrieval eligibility from ranking quality: first determine which content may be considered, then choose the most useful evidence from that authorized set.

An effective lab stores a small corpus with overlapping vocabulary but different tenant ownership. Ask the same question as users with different permissions and inspect the retrieved document identifiers, not merely the generated answer. Record whether a denied document could still appear in traces, caches or error logs. Where the design uses vector indexes, also measure how embedding updates are synchronized with document deletion and permission changes; an old embedding can outlive the business fact it describes.

Request units turn design choices into economics

Two applications that return the same answer may consume very different request units. A point read of a known item within the correct partition is not equivalent to a broad cross-partition query with sorting. Index policies also involve trade-offs: an index that speeds a heavily used query may increase write cost and storage. The important question is whether the workload is read-heavy, write-heavy, bursty, or dominated by expensive searches that the product team has not yet observed.

Set a repeatable benchmark with realistic document sizes, partition distributions and concurrency. Examine query metrics, throttling responses and regional latency while changing one design variable at a time. A 429 response is a capacity or demand signal to handle with appropriate retry behavior and investigation, not a reason to retry indefinitely. Check the SDK connection lifecycle as well: rebuilding a client for each request creates avoidable overhead that can masquerade as a database bottleneck.

Memory needs deliberate forgetting

An agent that never forgets is not necessarily a more capable agent. Conversation state may need short retention, profile preferences may require user-controlled correction, and event history may need distinct audit rules. Azure Cosmos DB time-to-live settings can help retire stored items, but applications should examine how generated summaries, vector representations and downstream copies are removed as well. Retention is a system behavior, not just one property on a container.

Test updates to conflicting memories. A customer changes their shipping address after a previous conversation and the assistant still quotes the old one. The correction path should identify authoritative records, invalidate stale context and preserve the provenance needed to explain the resulting answer. Optimistic concurrency and change-feed processing can support controlled updates, but neither guarantees correct business precedence unless the application defines which source wins.

Design the failure, not only the happy path

Cosmos DB’s replication, consistency choices, backup options and failover configuration shape how an application behaves during regional faults. A lower-consistency read may offer operational benefits but is not appropriate for every immediate read-after-write expectation. A replicated deletion may propagate quickly, so a replica is not a substitute for a tested restore process. Consider regional distribution, encryption, data-plane permissions and observability together rather than as independent checklist items.

For DP-420 practice, build a small assistant backend and deliberately create a hot partition, a stale memory, an unauthorized retrieval and a regional disruption scenario. Explain what evidence you would inspect and how you would change the design. The useful result is not merely knowing which Cosmos DB feature exists: it is being able to defend its behavior under real traffic, with predictable security, freshness and cost.

  • img