Microsoft AI-200: Identity and Secrets for AI Apps

AI-200 directly tests secure Azure solutions, including Key Vault secret rotation and retrieval and Azure App Configuration. The AI-200 exam from Microsoft also expects candidates to connect to databases, messaging services, Functions, containers, and monitoring systems, so identity becomes the connective tissue across the entire application.

The strongest default is to use workload identity and least privilege wherever Azure supports it, then use a secret store only for credentials that cannot be replaced by identity. Key Vault and managed identities provide the Azure-specific pattern, while cloud identity and least privilege provide the broader model for roles, service identities, and least privilege.

Security questions become easier when you separate four things: human identity, workload identity, configuration, and secrets. They have different owners, rotation needs, audit trails, and failure modes.

Separate human administrators from workload identities

applications should not run under developer accounts and developers should not need the application’s production credentials. In practice, that means human accounts, groups, privileged roles, managed identities, service principals, deployment identities, and break-glass administration. Assign runtime access to workload identities and administrative access to controlled human roles. Use separate identities across environments so a development workload cannot accidentally operate production data.

Shared identities erase accountability and make credential rotation or employee offboarding risky. Useful evidence includes distinct principals, role assignments, sign-in logs, access reviews, and production deployment records. AI-200 candidates should be comfortable connecting applications to Azure services without embedding personal credentials. Privileged deployment access should not be inherited by the running application.

Prefer managed identity when the target service supports it

For identity architecture, the key principle is that managed identity removes the need to distribute and rotate a client secret for many Azure-to-Azure calls. The working parts are system-assigned and user-assigned identities, RBAC, token acquisition, service endpoints, environment separation, and lifecycle. Choose the identity type based on ownership and reuse, then grant only the target roles required. User-assigned identities can simplify reuse across resources, while system-assigned identity lifecycle follows the resource.

A managed identity can still be over-privileged, so removing a secret does not automatically create least privilege. Validate the result with identity configuration, scoped roles, successful and denied calls, audit logs, and lifecycle behavior. Key Vault and managed identities shows the production pattern of pairing managed identities with protected secret/configuration access. The scope decision should be deliberate.

Use Key Vault for secrets that must exist

Identity design is easier to audit when you separate the design goal from the implementation detail: when a credential, key, or certificate cannot be replaced with workload identity, it should have a protected owner, retrieval path, rotation process, and audit trail. The implementation normally spans secrets, keys, certificates, vault access, managed identity retrieval, rotation, versioning, caching, and outage behavior. Keep secret material outside source code and images, retrieve it at runtime through a scoped identity, and design rotation before production. Test what happens when a secret version changes while instances are running.

A secret can be stored in Key Vault and still be unsafe if an application caches it indefinitely or the retrieval identity has excessive access. The evidence that matters is vault access logs, rotation tests, version history, secret inventory, expiry alerts, and failure behavior. Key Vault rotation and retrieval are explicitly named in the current AI-200 skills. Rotation is an application behavior as well as a vault setting.

Keep configuration separate from secret material

application configuration needs versioning and environment management, but most settings should not be treated as credentials. Operationally, the design touches App Configuration, feature flags, endpoints, model names, limits, connection metadata, environment overrides, and Key Vault references. Put non-secret settings in configuration services and reference protected secrets rather than copying them into configuration values. Configuration changes can materially change AI behavior even when code does not change, so they deserve review and observability.

Teams often overuse secret stores for ordinary settings or accidentally put credentials into environment-specific config files. A review should look for configuration history, environment separation, secret references, access policies, and deployment tests. Azure App Configuration is a direct AI-200 objective. Tag releases with configuration versions where practical.

Least privilege must follow the service graph

Evaluate this as an access-control design rather than a feature checklist: each backend component should receive the minimum data-plane and control-plane permissions required for its responsibility. The system includes Cosmos DB, PostgreSQL, Redis, Service Bus, Event Grid, Functions, storage, model services, Key Vault, and monitoring. Map every workload identity to the operations it performs and scope roles to the narrowest practical resource boundary. Different tasks may need read, write, listen, send, invoke, or management rights; do not collapse them into one permission set.

One broad contributor role can allow a compromised service to alter infrastructure or access unrelated data. Prove the design with role matrix, access tests, denied operations, Azure activity/audit logs, and periodic review. Identity scenarios often combine several services, so reasoning across the graph is more important than memorizing role names. Review permissions when architecture changes.

Protect CI/CD and deployment identities

deployment automation usually needs different privileges from the runtime application and can become a powerful attack path. From there, the engineer or manager has to coordinate federated credentials, service principals, GitHub or Azure DevOps identity, environment approvals, infrastructure deployment, application deployment, and secret access. Use short-lived or federated credentials where possible and scope the pipeline to the exact environment and deployment actions. Separate build from deploy and separate production approval from routine code execution.

A long-lived high-privilege pipeline secret can bypass all runtime least-privilege controls. The most persuasive evidence is federation configuration, pipeline permissions, approval logs, deployment history, and secret-free workflows. General CI/CD fundamentals controls are directly relevant to the development lifecycle around AI-200. Rollback rights should also be scoped and tested.

Identity failures should be diagnosable from evidence

authentication, token acquisition, role assignment, network access, and resource-level authorization can fail independently. In practice, that means identity availability, token audience, tenant, role scope, propagation, private endpoints, firewall, service logs, and request IDs. Trace the request from identity acquisition to target-service authorization rather than changing several permissions at once. cloud networking fundamentals can help separate network reachability from authorization when private connectivity is involved.

Teams often grant broader roles when the real problem is the wrong identity, wrong resource scope, network path, or stale configuration. Useful evidence includes token claims where appropriate, sign-in logs, Azure activity logs, service audit logs, denied-operation codes, and network traces. Troubleshooting is a full 20–25% domain together with security and monitoring. A diagnostic change should be reversible and narrower than a permanent fix.

Design secret rotation as a no-drama operational event

For credential operations, plan on the fact that rotation should be routine enough that the team can perform it without emergency downtime. The working parts are expiry, version overlap, consumer refresh, connection pools, certificates, staged rollout, rollback, alerting, and owner accountability. Choose a rotation process that lets old and new credentials overlap only as long as required and verify every consumer updates. Practice an emergency rotation for a suspected compromise, not only a scheduled rotation.

A secret with a rotation policy can still expire in production if an application never refreshes it or ownership is unclear. Validate the result with rotation runbooks, automated alerts, successful refresh, access logs, and removal of old versions. AI-200 explicitly names rotation, making operational handling important. That tests whether the identity and deployment model is actually recoverable.

Monitor identity and secret behavior as security telemetry

Identity monitoring improves when you separate the design goal from the implementation detail: identity architecture is incomplete if anomalous access, failed secret retrieval, privilege changes, and unexpected workload activity are invisible. The implementation normally spans sign-in logs, role changes, vault access, denied requests, secret age, unusual service access, alerting, and incident correlation. Define alerts that correspond to real response actions instead of collecting every event at the same severity. Include deployment identity events because a compromised pipeline can change runtime identity or configuration.

Security teams miss compromise when service identities are considered ‘non-human’ and excluded from behavioral monitoring. The evidence that matters is centralized logs, change events, access baselines, alert investigations, and incident timelines. OpenTelemetry and KQL support application troubleshooting, while Azure audit sources support security investigation. Monitoring should feed access review and risk reassessment.

Production designs should also include an access-review rhythm for workload identities. Service identities often outlive the feature or environment that created them, leaving stale permissions that are difficult to notice because no person signs in interactively. Review ownership, last use, assigned roles, and downstream dependencies before removing or narrowing access.

Break-glass access deserves the same discipline. Emergency credentials or privileged roles should be tightly controlled, monitored, tested, and excluded from routine automation. A recovery path is valuable only if the team can use it without turning a temporary emergency privilege into a permanent backdoor.

Document the emergency path too.

Finally, include identity changes in release testing. Moving an application to a new managed identity, narrowing a role, rotating a certificate, or changing a vault reference can fail even when the code artifact is unchanged. A safe rollout proves both allowed and denied operations, verifies telemetry for the new principal, and confirms that rollback does not restore stale credentials. That operational discipline is central to AI-200 because identity and secret handling must survive normal deployment, not only initial configuration.

  • img