Amazon AWS MLA-C01 to MLA-C02: Model Development
MLA-C02 changes the AWS Machine Learning Engineer – Associate model-development domain more visibly than it changes data preparation. MLA-C01 devoted 26% of scored content to ML Model Development: choosing a modeling approach, training and refining models, and analyzing performance. MLA-C02 reduces the domain slightly to 24% but expands it into “ML Model and Foundation Model Development,” adding foundation-model selection, RAG architecture, customization, prompt engineering, Bedrock evaluation, GenAI metrics, and human evaluation.
Because MLA-C01 English testing ended September 28, 2026, candidates should treat it as legacy context rather than a current English exam. The Amazon AWS MLA-C01 material preserves the historical scope, while AWS AI certifications and MLA-C02 preparation reflect the active direction.
MLA-C01 expected candidates to match algorithms and model families to business problems while considering data, complexity, interpretability, latency, and performance. Those fundamentals remain in MLA-C02. Candidates still need to distinguish classification, regression, clustering, forecasting, anomaly detection, recommendation, and other common problem types and understand when a managed AI service may be more appropriate than a custom model.
The decision should begin with the requirement, not the service. What output is needed? How much labeled data exists? How important is interpretability? What latency and cost are acceptable? How often will the model change? A simpler model that meets the requirement may be preferable to a more complex approach that is harder to explain, operate, or retrain.
Foundation models introduce a different selection problem. Candidates need to consider modality, quality, latency, context length, cost, language support, safety characteristics, and customization options. Amazon Bedrock exposes multiple model families, so “use a foundation model” is not a complete answer. The selected model should fit the task and operational constraints.
Evaluation should happen before deep integration. A small representative test set can reveal whether a model follows instructions, retrieves the right knowledge, produces acceptable latency, and stays within cost targets. Models should be compared on the application’s real success criteria rather than public benchmark reputation alone. The best model for summarization may not be best for extraction, coding, tool use, or multilingual support.
MLA-C02 explicitly expects candidates to select RAG patterns. RAG can improve groundedness and freshness by retrieving external knowledge at inference time instead of trying to encode all domain information into model weights. The design includes document preparation, embeddings, vector retrieval, filtering, ranking, prompt assembly, model choice, and evaluation.
RAG is not automatically the correct answer. If the model needs to learn a consistent style or task behavior, fine-tuning may fit better. If information is small and static, a simpler prompt/context approach may be sufficient. If retrieval quality is poor, adding a larger model may not solve the problem. Candidates should know which part of the system is failing before choosing a customization technique.
MLA-C01 covered SageMaker built-in algorithms, common frameworks, hyperparameters, automatic model tuning, early stopping, distributed training, overfitting, underfitting, regularization, and performance analysis. Those remain relevant. MLA-C02 still expects candidates to train models, tune parameters, manage experiments, and assess performance.
Hyperparameter tuning should have a defined objective and search space. More trials are not automatically better if the metric does not represent business success. Training should also be reproducible: code, data version, features, parameters, environment, and model artifact need traceability. Reproducibility is what allows teams to distinguish a true improvement from a lucky run or hidden data change.
Foundation-model customization can include prompt engineering, fine-tuning, continued pre-training, and other model-specific approaches. MLA-C02 expects candidates to recognize when those techniques are appropriate and what trade-offs they introduce. Fine-tuning can improve specialized behavior but requires quality data, evaluation, cost, governance, and lifecycle management.
Customization should follow a problem diagnosis. If the issue is missing current knowledge, RAG may be better. If the issue is task instruction, prompting or examples may be enough. If the issue is domain behavior that prompts cannot reliably produce, fine-tuning may be justified. The strongest answer is the least complex technique that meets quality, latency, cost, and governance requirements.
MLA-C02 adds AI-specific customization and prompt-management concepts. Prompts should be treated as versioned application assets rather than ad hoc text. Instructions, context, examples, constraints, output schema, tool descriptions, and safety boundaries can materially change behavior. Teams should test changes against representative scenarios before promoting them.
Prompt evaluation should include task success and failure behavior. A prompt that works on easy examples may fail under ambiguous input, long context, conflicting instructions, or adversarial text. Versioning enables rollback and comparison. This makes prompt development look increasingly like software engineering: controlled changes, tests, metrics, and deployment discipline.
Traditional ML evaluation still includes metrics such as accuracy, precision, recall, F1, AUC, RMSE, or task-specific measures. MLA-C02 adds AI evaluation concepts such as BLEU, ROUGE, BERTScore, semantic similarity, output quality assessment, bias detection, model-as-judge patterns, and human evaluation. The candidate should know that no single metric is universally sufficient.
Metric choice follows the cost of error. In fraud detection, false negatives and false positives have different consequences. In generated answers, groundedness, relevance, completeness, harmfulness, and format compliance may all matter. Business metrics such as user success, time saved, or escalation rate can be more important than a laboratory score. A robust evaluation plan combines technical and outcome measures.
Some generated outputs cannot be evaluated reliably by automated metrics alone. Human reviewers can assess usefulness, factuality, tone, safety, or domain correctness. MLA-C02 explicitly includes integrated human evaluation frameworks. The challenge is making human evaluation consistent enough to be useful.
Reviewers need clear rubrics and representative examples. Inter-rater disagreement should be measured rather than ignored. Sensitive or high-impact use cases may require expert reviewers and escalation criteria. Human feedback can also become training or preference data, which means the evaluation process itself becomes part of the data-governance lifecycle.
The arrival of foundation models does not remove classic ML failure modes. Training performance that is much better than validation performance can indicate overfitting. Poor performance on both may indicate underfitting, weak features, insufficient capacity, or a mismatched model. Learning curves, validation metrics, feature importance, error analysis, and regularization can help identify the cause.
For GenAI systems, an analogous diagnostic mindset still applies. A weak answer may come from the base model, prompt, retrieval, data, context construction, tool output, or evaluation method. Changing the foundation model without locating the failure can waste cost and hide the root problem. Candidates should learn to isolate the component that limits quality.
MLA-C02 explicitly includes reproducible experiments using tools such as MLflow on SageMaker AI and Bedrock evaluation or prompt-management capabilities. An experiment record should identify data, code, model, parameters, environment, metrics, and artifacts. This allows teams to compare variants and reproduce a result after the original notebook or developer session is gone.
Tracking also supports governance. A deployed model should be traceable to the experiment and approval that produced it. If a new model performs worse, teams need a previous known-good version. If a regulator or internal reviewer asks why a model changed, experiment history provides evidence rather than relying on memory.
Model development is not finished when a metric is maximized. Training cost, inference latency, throughput, memory, endpoint size, context length, token usage, and operational complexity affect whether the model is usable. A more accurate model may be a poor production choice if it is too slow or too expensive for the workload.
MLA-C02 emphasizes trade-offs for both traditional ML and foundation models. Engineers should compare model size, serving infrastructure, batching, quantization or optimization options, caching, retrieval design, and workload shape. The correct decision is the model system that meets the business requirement reliably, not the model with the largest benchmark number.
The MLA-C01 to MLA-C02 transition covers the broader exam change. For Domain 2, retain your understanding of modeling approaches, SageMaker training, hyperparameter tuning, overfitting, evaluation, and experiments. Then add foundation-model selection, RAG, AI customization, prompt engineering, GenAI evaluation, human review, and AI-specific cost/latency trade-offs.
That transition is larger than a simple service-name update, but it builds on familiar engineering logic. Define the requirement, choose the simplest appropriate approach, train or customize deliberately, measure with the right metrics, keep experiments reproducible, and evaluate production constraints. MLA-C02 expands the model types; it does not remove the discipline that made MLA-C01 model development useful.
A useful transition exercise is to take one familiar MLA-C01 modeling scenario and solve it twice. First, solve it as a traditional ML problem with feature preparation, model selection, training, tuning, and validation. Then ask whether MLA-C02 introduces a credible FM or RAG alternative, what new quality metrics would be required, and whether the added latency, cost, safety, or governance burden is justified. This prevents the new GenAI material from becoming an automatic answer to every modeling problem. MLA-C02 expands the solution space; it still rewards engineers who choose the simplest approach that meets the requirement.
Also practice identifying the cheapest experiment that can disprove a modeling idea early. A small baseline, retrieval test, prompt evaluation, or shadow comparison can prevent expensive tuning of the wrong approach and is exactly the kind of engineering judgment the updated exam rewards. For model-development scenarios, compare candidate approaches on data fit, evaluation evidence, latency, cost, operational complexity, and the failure mode that matters most to the business requirement.
