Amazon AWS MLA-C01 to MLA-C02: Data Preparation

AWS changed the Machine Learning Engineer – Associate exam at the end of September 2026, but the data-preparation skills built for MLA-C01 did not disappear. MLA-C01 English testing ended on September 28, and the MLA-C02 beta became the current English exam on September 29. Domain 1 remains 28% of scored content, which signals strong continuity, but the scope expands from “Data Preparation for Machine Learning” to “Data Preparation for ML and AI.”

MLA-C01 is now transition context rather than current preparation. The legacy Amazon AWS MLA-C01 material preserves the historical scope, while AWS AI certifications and MLA-C02 preparation reflect the active direction. Preparation should preserve durable ML data skills while identifying the new AI and generative-AI work introduced by MLA-C02.

The 28% domain weight hides a meaningful scope expansion

MLA-C01 Domain 1 already expected candidates to ingest and store data, transform and engineer features, validate data quality, address bias, and prepare datasets for modeling. MLA-C02 keeps those tasks but expands the data types and downstream use cases. Candidates now need to think about data for traditional ML and AI systems, including vector stores, embeddings, text and image data, RAG pipelines, foundation-model customization, and AI-specific integrity concerns.

The right transition strategy is not to discard MLA-C01 notes. Instead, label each skill as retained, expanded, or new. Storage decisions, ingestion, data formats, transformation, data quality, and bias remain core. Vector databases, document chunking, embedding preparation, prompt-response data, multimodal data, and FM training or distillation preparation are the areas where legacy preparation needs deliberate expansion.

Storage selection still begins with access patterns and economics

MLA-C01 expected candidates to understand how sources such as Amazon S3, EBS, EFS, RDS, DynamoDB, streaming platforms, and distributed processing environments fit data ingestion. MLA-C02 retains that foundation. Storage decisions still depend on data structure, access pattern, latency, throughput, scale, cost, durability, compliance, and how the data will be consumed by downstream training or inference workflows.

The AI expansion adds more diverse assets. Text corpora, documents, images, audio, embeddings, and vector indexes may have different storage and retrieval needs from tabular training data. Candidates should be able to separate raw source storage from feature or vector-serving layers. A vector database is not a replacement for every source system; it is an optimized retrieval layer for similarity-based use cases.

Data formats matter because they change cost and pipeline behavior

Choosing CSV, JSON, Parquet, ORC, Avro, or another format affects storage size, schema behavior, scan efficiency, interoperability, and processing cost. Columnar formats can reduce the amount of data read for analytical or feature-engineering workloads. Row-oriented or semi-structured formats may be convenient for interchange but more expensive for repeated scans. The correct format depends on how the pipeline reads and transforms the data.

MLA-C02 also expects candidates to consider diverse AI data types. Images and audio are not handled like tabular records, and document pipelines may preserve both original content and extracted metadata. The candidate should ask whether downstream systems require random access, sequential processing, batch scans, streaming updates, or similarity search. Format and storage decisions should be made together rather than independently.

Feature engineering remains central for traditional ML

Scaling, normalization, binning, log transforms, encoding, feature splitting, and feature selection remain core ML preparation techniques. SageMaker Data Wrangler, Feature Store, AWS Glue, Spark, and related services continue to support these workflows. The purpose is to make features consistent, informative, and suitable for the chosen model rather than simply applying transformations because a tool offers them.

Reproducibility matters. Training and inference should apply compatible transformations, and feature definitions should be versioned or governed so a model is not trained on one representation and served another. Leakage is another key risk: features that contain future information or target-derived data can produce impressive validation results and poor real-world performance. A strong candidate recognizes preparation mistakes that invalidate the experiment.

Data quality and bias controls carry directly into MLA-C02

Missing values, duplicates, outliers, inconsistent labels, schema drift, corrupt records, and class imbalance can weaken any ML system. MLA-C01 already expected candidates to validate quality and manage bias using services and techniques such as AWS Glue Data Quality, DataBrew, SageMaker Clarify, splitting, shuffling, augmentation, or resampling. Those skills remain useful in MLA-C02.

AI systems increase the importance of content quality because training or grounding data can influence generated behavior. Bias can appear in text, images, or multimodal data, not only structured labels. Candidates should think about representativeness, harmful or restricted content, provenance, duplication, and whether the data supports the intended use. “More data” is not automatically “better data.”

Vector data and embeddings are a major MLA-C02 addition

MLA-C02 explicitly adds scalable vector databases and embedding-based data preparation. Embeddings transform text, images, or other content into numerical vectors that can be compared for semantic similarity. Preparing this layer involves choosing an embedding model, deciding what content is embedded, normalizing metadata, selecting a vector store, and designing update and deletion behavior.

The storage decision must reflect expected query volume, latency, filtering, scale, security, and consistency. Amazon OpenSearch Service, relational systems with vector extensions, and other AWS-supported patterns can serve different needs. The exam does not require candidates to invent vector algorithms, but it does expect them to understand why an AI application needs a vector index and what operational trade-offs come with it.

RAG preparation turns documents into retrievable evidence

Retrieval Augmented Generation requires more than copying files into a bucket. Documents often need extraction, cleaning, chunking, metadata, embeddings, and indexing. Chunk size affects retrieval precision and context completeness. Metadata enables filtering and governance. Source identifiers allow applications to trace retrieved information back to an authoritative document. Poor preparation can make a strong foundation model produce weak grounded answers.

MLA-C02 explicitly adds RAG document preparation, including chunking strategies and metadata extraction. Candidates should reason from the application: what unit of information should be retrieved, how frequently does the source change, how are permissions represented, and how will removed or corrected content be reflected in the index? Data lifecycle becomes part of answer quality.

Fine-tuning and distillation require purpose-built training data

Traditional supervised learning already depends on well-labeled datasets. MLA-C02 extends preparation into foundation-model fine-tuning, continuous pre-training, and model distillation. These workflows may require prompt-response pairs, instruction data, domain text, preference data, or teacher-model outputs. The data must match the target behavior and be screened for quality, privacy, safety, and rights to use.

More examples are not useful if they encode inconsistent instructions or undesirable behavior. Teams should define expected output style, task boundaries, sensitive content handling, and evaluation criteria before assembling the dataset. Train/validation separation also remains important; using evaluation examples in the tuning data makes performance appear better than it is.

Security and compliance should be applied before modeling

ML data can contain personal, confidential, regulated, or proprietary information. MLA-C01 already included classification, anonymization, masking, encryption, residency, PII, and PHI concepts. MLA-C02 retains and expands that responsibility as AI systems consume more unstructured and multimodal data. Privacy controls are cheaper and safer when built into ingestion and preparation rather than applied after a dataset has been widely copied.

Access should follow least privilege, and pipelines should record enough lineage to explain where training or grounding data came from. Encryption in transit and at rest is necessary but not sufficient. Teams also need policies for retention, deletion, approved use, and downstream derivatives such as features, embeddings, or model artifacts. AI systems can reproduce sensitive information if governance fails upstream.

Preparation is an operational workflow, not a one-time notebook task. Production pipelines should be repeatable, idempotent where practical, monitored, and able to recover from partial failure. Engineers need to know which source version produced a training set, which transformation code ran, which records were rejected, and what quality checks passed. Without that evidence, model debugging becomes guesswork.

Streaming workflows add additional concerns such as ordering, late data, backpressure, checkpointing, and schema evolution. Batch workflows need reliable partitioning, retries, and cost controls. The exact AWS service can vary. The design principle is stable: data preparation should produce reproducible assets with clear lineage and measurable quality.

The MLA-C01 to MLA-C02 transition explains the broader certification change. For Domain 1, the efficient study plan is more focused. Keep MLA-C01 strength in storage, ingestion, formats, transformation, feature engineering, bias, data quality, security, and compliance. Then add vector data, embeddings, multimodal preparation, RAG documents, and FM customization data.

This is a favorable transition for candidates with solid ML engineering fundamentals. The core discipline is unchanged: understand the data before training a model. MLA-C02 simply extends that discipline into newer AI systems and data representations. The candidate who can explain how data quality, lineage, security, retrieval, and transformation affect model behavior will be better prepared than someone who memorizes a new list of services without understanding the pipeline.

One final transition habit is to separate data-preparation tools from data-preparation outcomes. MLA-C02 may name newer AI services or features, but the engineer is still responsible for producing data that is accessible, representative, secure, traceable, and appropriate for the modeling task. When a scenario offers several AWS services, first identify the required outcome—stream ingestion, feature reuse, vector retrieval, document preparation, quality validation, or bias mitigation—then choose the service that fits the operational constraints. That reasoning is more durable than memorizing the latest console label.

For study review, be able to explain why a chosen data path is appropriate and what quality check would fail it. That one habit connects storage, preprocessing, bias, RAG preparation, and governance without relying on service-name recall.

  • img