Amazon AWS MLA-C01 to MLA-C02: MLOps and Operations

MLOps is the part of machine learning engineering that begins after a model appears to work. It covers how releases are versioned, monitored, secured, scaled, evaluated, retrained, and recovered when production behavior changes. The September 2026 AWS exam transition makes that scope broader: English MLA-C01 testing ended September 28 and MLA-C02 beta began September 29, expanding “monitoring, maintenance, and security” into operation of both ML and AI solutions.

MLA-C01 now provides legacy context for English-language candidates. AWS AI certifications place MLA-C02 in the active branch. The transition preserves classic MLOps while adding foundation-model and agent-workflow operating concerns.

Production ownership begins with an observable release

A model release should be traceable to source code, training data or data references, feature logic, model artifact, evaluation result, deployment configuration, approvals, and the pipeline that promoted it. Without that chain, a production incident becomes guesswork. Teams may know which endpoint is failing without knowing which data or code produced the model behind it.

Version the artifacts that matter and expose that version in operational metadata. A rollback decision should identify the previous known-good model and its compatible environment, not merely an older timestamp. Reproducibility is a reliability control.

Monitor service health separately from model quality

Endpoint latency, errors, CPU or GPU utilization, memory, queue depth, throttling, and availability describe service health. Prediction accuracy, ranking quality, calibration, business conversion, or human evaluation describe model behavior. One can look healthy while the other degrades.

Build separate dashboards and alarms for both layers. A low-latency endpoint producing systematically poor predictions is not healthy. A high-quality model that times out during peak traffic is not healthy either. Production acceptance needs both operational and analytical evidence.

Data quality monitoring catches problems before drift metrics do

Missing columns, changed categories, invalid ranges, schema changes, stale features, duplication, and broken pipelines can distort inference immediately. These are often data-quality failures rather than gradual statistical drift. Monitor input completeness, freshness, schema, distribution, and feature-generation success before relying on higher-level drift indicators.

When a model suddenly changes behavior, first ask whether the input represents the same thing it represented during training. A distribution shift may be real market change, but it can also be a pipeline bug that remapped a field. Context prevents teams from retraining a model to compensate for broken data.

Drift needs a response policy, not just a chart

Data drift, concept drift, and changing business conditions can reduce model effectiveness. A drift signal should therefore connect to a decision: investigate, compare performance against labeled outcomes, retrain, change thresholds, or accept the shift because the model remains useful.

Set thresholds from risk and observed behavior rather than arbitrary percentages. High-risk models may require tighter monitoring and human review. Low-impact models can tolerate more variation. Record which drift signals are advisory and which trigger workflow actions so automation does not retrain or redeploy blindly.

Retraining should be governed like any other release

Automated retraining can keep a model current, but a new training run is not automatically an improvement. The pipeline should validate data, compare metrics with the incumbent model, check fairness or other required constraints, record lineage, and apply deployment gates. A schedule or drift event can trigger evaluation without automatically granting production approval.

Define rollback for retrained models and retain the evidence behind promotion. If labels arrive slowly, business metrics or human review may be needed before full rollout. Production MLOps balances freshness with release safety.

Cost is an operational metric

ML systems can remain technically healthy while becoming economically unhealthy. Monitor endpoint hours, accelerators, storage, data transfer, feature processing, pipeline runs, model-training jobs, and idle provisioned capacity. Tie cost to workload or model version where possible so teams can see what changed.

Optimization should respect service objectives. Rightsizing an endpoint can reduce cost but increase tail latency. Scaling to zero can save money but introduce cold starts. Batch processing can be cheaper but violate freshness requirements. MLOps treats cost as one dimension of the production design rather than an afterthought.

Security controls belong inside the ML lifecycle

Production workflows touch data stores, code repositories, registries, container images, endpoints, secrets, pipelines, and monitoring systems. Use least-privileged roles, encryption, network boundaries, image scanning, artifact controls, and auditable administrative actions. Model files and data are sensitive assets even when they do not look like traditional source code.

Separate roles for training, approval, deployment, and monitoring where the risk justifies it. Protect emergency operations as well: a responder should be able to roll back or isolate a failing service without receiving permanent broad privilege.

CI/CD and ML pipelines solve different parts of the lifecycle

CI/CD manages versioned software and infrastructure changes. ML pipelines manage data preparation, training, evaluation, registration, and sometimes deployment. Mature environments connect them while preserving the difference. A code change may trigger training; a new approved model may trigger deployment without a code change.

Design triggers and gates so teams know why a run started and what artifact advanced. Avoid pipelines where every commit automatically trains and deploys an expensive model regardless of impact. Automation should encode policy, not eliminate judgment.

Incident response needs model-specific evidence

A production incident may involve bad predictions, poisoning, unauthorized endpoint access, broken feature data, malicious input, model extraction attempts, or simple infrastructure failure. Preserve model version, request samples where policy allows, feature or retrieval context, deployment events, IAM activity, and monitoring history.

Response plans should distinguish containment options. Disabling an endpoint, switching traffic to a previous model, blocking a data source, revoking credentials, or falling back to rules have different business effects. Predefined decision paths make model incidents easier to manage under pressure.

Foundation-model systems create metrics that classic predictive models do not. Token usage, context length, refusal rate, groundedness, citation quality, safety outcomes, tool errors, and prompt or retrieval versions can affect service quality. There may be no single accuracy number that captures the user experience.

Define evaluation sets and operational quality indicators that match the task. For a summarization system, factuality and omission matter. For a support agent, resolution rate, unsafe actions, tool success, and escalation behavior may matter more. Monitor the system around the model, not only the model call.

Retrieval-augmented generation depends on source documents, parsing, chunking, embeddings, indexes, permissions, retrieval filters, and the model. A stale or partially updated index can make the model appear unreliable even when inference is unchanged. Monitor freshness and ingestion failures as production quality signals.

When answers degrade, separate retrieval failure from generation failure. Inspect whether relevant chunks were available and selected before changing prompts or models. That layered diagnosis is the GenAI equivalent of checking input features before retraining a traditional model. Agentic workflows add action safety to operations.

An agent can call tools and change external systems, so operations must include tool permissions, action logging, loop detection, timeout, retry, approval boundaries, and compensation when a multi-step action partially fails. Model quality monitoring alone cannot tell you whether the workflow is safe.

Track tool success, repeated calls, unexpected branches, human escalation, and irreversible actions. Restrict tools to the smallest useful capability. Production AI becomes an operational system with privileges, state, and side effects; MLOps must evolve accordingly. The transition preserves the core operating discipline.

MLA-C02 preparation and the MLA-C01 to MLA-C02 transition clarify which operating skills remain durable and which AI responsibilities are new. The durable lesson is that production ML and AI must be versioned, observable, secure, cost-aware, and recoverable.

MLA-C02 adds FM, RAG, and agent operating concerns, but it does not replace classic MLOps. Candidates should carry forward release control, data/model monitoring, retraining governance, security, cost management, and incident response, then extend the same discipline to the additional components of modern AI solutions. Model registries should express promotion state.

A registry is more useful when it distinguishes candidate, approved, deployed, deprecated, and rejected model versions rather than acting as a file shelf. Attach lineage, evaluation metrics, approval evidence, framework/runtime requirements, and intended environment. Promotion should be a controlled state change that a pipeline can enforce and an auditor can reconstruct.

For foundation-model or prompt-centric applications, the equivalent inventory may also need prompt versions, retrieval configuration, guardrail versions, and evaluation results. The governed production artifact is increasingly a bundle of components rather than one model file. Service quotas are part of production capacity.

Training jobs, endpoints, accelerators, API rates, foundation-model throughput, and supporting services all have quotas or practical limits. A workload can be correctly autoscaled and still fail because an account or Region cannot allocate more capacity. Capacity planning should identify quota-dependent bottlenecks before launch and monitor consumption as demand grows.

Include quota increase lead time and regional availability in recovery planning. A failover design that assumes identical accelerator or FM capacity in another Region may not meet its objective unless that capacity has been validated in advance. Runbooks should name both technical and model-quality actions.

An alarm should point responders toward a decision, not only a dashboard. A latency incident may require scaling or rollback; a quality incident may require disabling a model version, reverting retrieval data, or increasing human review. Document owners, evidence, containment options, rollback steps, and criteria for returning the system to normal operation.

Data and model retention policies also belong in the operating model. Keeping every training snapshot, feature set, prompt trace, and model indefinitely can create privacy, security, and cost problems, while deleting them too quickly can make incidents or reproducibility impossible. Define retention from regulatory, investigative, rollback, and learning needs, and ensure deletion propagates to derived artifacts where required.

Production review should periodically challenge whether each monitored signal still predicts useful action. Dashboards accumulate metrics easily; mature MLOps removes or redesigns indicators that no longer influence decisions and strengthens the few that reliably expose risk.

  • img