Microsoft AI-300: Hands-On MLOps and GenAIOps Practice

Microsoft recommends hands-on experience for AI-300, and the current blueprint contains many skills that are difficult to learn by reading alone: workspace assets, endpoints, rollout/rollback, drift monitoring, CI/CD, Foundry observability, evaluation, RAG optimization, and fine-tuned model operations. A useful lab does not need an expensive production environment. It needs enough scope to create observable lifecycle events and force you to diagnose them.

Lab one: build the Azure ML workspace baseline

Create or study an Azure Machine Learning workspace with a datastore, compute target, identity/RBAC, and a few managed assets. Record which identity performs each operation and where artifacts live. The lab should make the relationship between workspace, compute, data, environment, component, model, and endpoint tangible.

Add a private-networking thought exercise. Even if the full network is too costly to build, draw the DNS, subnet, private endpoint, identity, storage, Foundry/Azure ML, and monitoring relationships. Then choose one blocked connection and explain the evidence you would use to isolate it.

Add one “ownership matrix” to the combined lab: data scientist, ML engineer, platform/DevOps engineer, application engineer, and security/governance stakeholder. Assign who owns training data, pipeline, endpoint, Foundry app, evaluation baseline, monitoring, incident response, and rollback. The role handoffs make operational gaps visible.

Lab two: deploy a model with safe rollout and rollback

Deploy a simple model to a real-time or batch endpoint. Introduce a second version and design a progressive rollout. Test the endpoint, capture latency/error evidence, and define the signal that would trigger rollback. A deployment exercise without rollback criteria misses a major production concern.

Add a rollback drill after every deployment lab. Do not wait for the final combined scenario. Record which model/app version is known good, which infrastructure changes are reversible, and which data/state changes need special care. Repeated rollback practice makes production safety a habit.

When the lab is complete, review it as a production service rather than a demo. Ask whether another engineer could deploy it, monitor it, evaluate it, roll it back, and understand its cost and safety controls without relying on your memory.

Lab three: monitor drift and retraining triggers. Create or simulate a change in input data and observe how drift or model-performance monitoring should respond. Define thresholds, alerts, and the conditions that justify retraining. Then separate “data changed” from “model performance actually degraded.” Operations needs evidence before spending resources on retraining.

For drift, simulate both benign and harmful change. A new data source may shift distributions without reducing model performance, while another shift can break accuracy. Define what evidence distinguishes those cases and when retraining should be triggered automatically versus reviewed by a human.

Lab four: automate infrastructure and delivery

Use Bicep or Azure CLI concepts plus a GitHub Actions-style workflow to represent environment deployment and application/model change. Bicep and infrastructure as code make environment changes repeatable and reviewable. Keep secrets out of code, use managed identities where appropriate, and make the automation produce logs that show what changed.

Lab five: create a Foundry project with controlled access. Model the Foundry resource/project, RBAC, managed identity, network access, and dependent services. Even if you cannot reproduce every private-network configuration in a lab, draw the intended trust boundaries and explain which connection would break if a role or network rule were wrong.

Lab six: trace a generative AI request

Instrument a simple generative AI application or agent with logging/tracing. Capture latency, response time, token consumption, tool/retrieval steps, and errors. Reproduce one failure and find where it occurs. This trains observability as a debugging path rather than a dashboard decoration.

For GenAI tracing, include at least one agent or tool call. Capture which tool was selected, inputs/outputs, latency, error, and final response. If the answer is wrong, determine whether the model chose the wrong tool, the tool returned poor data, retrieval was weak, or the final generation misused correct context.

Lab seven: build a repeatable evaluation set. Create representative prompts/tasks with expected criteria and compare two versions of the application or prompt. prompt and model evaluation provides deeper evaluation context. Measure more than “looks good”: include relevance, groundedness, task success, safety, latency, or cost as appropriate.

For evaluation, create a small fixed regression set and rerun it after every meaningful prompt, retrieval, or model change. This builds the habit of comparing versions on the same tasks. Add one adversarial or edge case and one safety requirement so “average quality” does not hide a serious failure.

Lab eight: tune a small RAG system

Change chunk size, similarity threshold, retrieval strategy, embedding choice, or hybrid search configuration one variable at a time. Measure retrieval relevance and answer quality after each change. If quality changes, identify whether the improvement came from retrieval or generation rather than assuming every better answer proves the model changed.

For RAG, include source freshness. Replace or remove one grounding document and observe how retrieval/evaluation should detect the changed answer. Production RAG is partly a content lifecycle system, so operators need to understand how indexing and source updates can change behavior without application-code changes.

A final demo review should answer four questions without opening documentation: what is deployed, how it is identified/versioned, how quality and health are measured, and how you recover from a bad change. If any answer is vague, repeat the relevant lab before exam day.

Keep lab costs controlled by deleting idle resources and using small datasets/models. Efficient cleanup is part of good cloud operations and prevents the study environment from becoming a distraction.

Lab nine: treat fine-tuning as a lifecycle

Use a small synthetic or curated dataset to model a fine-tuning workflow. Track dataset version, training configuration, model version, evaluation, deployment, monitoring, and rollback. The operational lesson is that a fine-tuned model creates new artifacts and risks that need the same lifecycle discipline as any production model.

Add a lab around asset promotion. Register or manage a model, environment, or component in one workspace and design how it would be promoted or shared to another environment through a registry or controlled pipeline. The learning objective is reproducibility: production should not depend on a local notebook state that cannot be reconstructed by another team.

Add a lab around endpoint troubleshooting. Create a deliberately unhealthy deployment through a bad environment, missing dependency, identity issue, or configuration error in a safe setup. Use deployment logs and endpoint state to locate the problem. Then recover without deleting all evidence. Troubleshooting is more useful when you can explain the failed stage.

Add a lab around model archiving or retirement. Deploy two versions, deprecate one, and document which applications still depend on it before removal. Lifecycle management includes ending support safely, not only launching new versions.

Add a lab around registry use if available. Publish or reference a shared environment, component, or model and then consume it from another context. Change the version deliberately and inspect how the downstream workflow should control adoption. Shared assets become dependencies that deserve review.

Lab ten: run one combined failure exercise

For Machine Learning Operations Engineer Associate readiness, simulate an application that has higher latency, worse quality, and rising token cost after a deployment. Use logs, traces, evaluation, retrieval metrics, and deployment history to isolate the cause. Responsible AI controls add governance requirements for safety, accountability, and review alongside the technical signals. Finish by proposing a rollback or optimization with evidence.

For cost, capture tokens or resource usage for several requests and compare them with measured quality. Try one optimization such as shorter context, different retrieval, caching, or model choice, then confirm the quality/cost trade-off. Cost reduction without quality evidence is incomplete optimization.

Add a CI/CD failure exercise. Break a pipeline variable, artifact path, permission, or deployment target and trace the failure from source control through automation logs to the cloud resource. The key skill is knowing which stage failed rather than restarting the whole pipeline until it works.

At the end, write a runbook for one production incident: rising errors on an endpoint or agent. Include deployment history, logs/traces, model/app metrics, rollback, user impact, and follow-up evaluation. A runbook proves that your lab skills can be turned into repeatable operations.

Add an observability acceptance test. Before calling a lab complete, confirm the dashboards or logs can answer: what version is running, how much traffic it receives, what latency/errors look like, and whether quality/cost changed. If those questions cannot be answered, the system is not truly operationalized.

End with a written handoff from “data scientist” to “operations engineer.” Include artifact versions, deployment method, evaluation baseline, monitoring thresholds, rollback, and owners. This exercise reflects the cross-team role Microsoft describes and turns lab work into an operational process.

Finally, tear down or archive the lab safely. Remove disposable endpoints, revoke temporary access, preserve useful evaluation/run records, and document anything intentionally retained. Cleanup is part of production discipline and proves you understand the lifecycle after experimentation is finished.

That operational handoff is the final proof that the lab taught more than feature navigation.

Clean up the environment and preserve only the evidence needed for review.

  • img