RAG and Grounding Pipelines for AI-103

AI-103 treats retrieval and grounding as core engineering skills rather than optional add-ons. The exam blueprint expects candidates to ingest and index content, configure semantic, hybrid, and vector retrieval, enrich documents and media, build RAG ingestion flows, and connect the results to applications and agent tools.

The general architecture in embeddings, vector databases, and RAG is useful background. For AI-103, focus on how Microsoft Foundry and Azure AI Search turn source content into grounded evidence that models and agents can use.

Start with source authority and access rules

Before indexing, identify which repositories are authoritative, which versions are current, and who may access the content. Retrieval quality is not enough if the index mixes obsolete policy with current policy or exposes restricted documents to the wrong user.

Keep source IDs, timestamps, ownership, and access metadata through the pipeline. Those attributes support filtering, deletion, freshness, and provenance later.

Ingest more than text when the use case requires it

AI-103 explicitly includes documents, images, audio, and video in retrieval and grounding pipelines. Different media need different extraction or representation steps before they can be indexed.

Do not convert everything into unstructured text if important layout, visual, or temporal information would be lost. Content Understanding and other Foundry tools can produce structured representations for downstream search and reasoning.

Use OCR and layout analysis when documents depend on structure

Scanned documents need OCR before text retrieval can work. Complex forms, tables, and multi-column layouts may need additional structure extraction so related fields remain connected.

Validate extraction before blaming the search layer. If a heading or table row is reconstructed incorrectly, no ranking method can retrieve the intended meaning reliably.

Choose semantic, vector, and hybrid retrieval deliberately

Semantic or lexical retrieval can preserve exact terms and document structure, while vector retrieval finds conceptually similar content. Hybrid approaches combine both signals and are often useful when users mix natural language with product names, codes, or domain terminology.

Evaluate on representative questions. The best retrieval method is the one that returns the authoritative evidence near the top for the workload, not the one with the most modern label.

Enrichment should improve searchability, not add noise

Built-in or custom skills can extract entities, tags, layout signals, captions, or other fields that help search. Each enrichment step should have a reason: filtering, ranking, routing, source attribution, or downstream reasoning.

Unnecessary enrichment increases processing time and can create misleading metadata. Test whether the added field improves retrieval or operations before making it a permanent pipeline stage.

Separate indexing quality from retrieval quality

An index can be healthy operationally while still producing weak search results. Monitor ingestion failures, document counts, vector generation, and index freshness separately from relevance metrics.

AI-103 includes both data-ingestion quality and index-health monitoring. Operational health tells you whether the pipeline is running; relevance evaluation tells you whether it is useful.

Assemble context around the user’s actual question

A RAG flow should select enough evidence to answer the request without filling the model context with unrelated material. Preserve source boundaries and useful metadata.

Long context is not a replacement for retrieval discipline. Irrelevant passages can distract the model or introduce contradictory statements even when everything fits inside the context window.

Grounded answers should retain source evidence

When the interface supports it, keep citations or source identifiers close to the claims they support. This lets users and downstream systems verify the answer and makes troubleshooting easier when two sources disagree.

Grounding quality is partly a product-design question: the application should make insufficient evidence visible rather than forcing the model to produce an answer every time.

Connect retrieval directly to agent tools when appropriate

Foundry agents can use search as a tool inside a larger workflow. The agent may decide when retrieval is needed, then combine the evidence with another tool or action.

Keep retrieval permissions aligned with the agent’s identity. A tool should not turn a broadly capable agent into a bypass around document access controls.

Evaluate RAG by stage

Use the approach in AI evaluation fundamentals to measure source extraction, retrieval recall, ranking, grounding, and final answer quality separately. When the response is wrong, you should know which stage failed.

Include queries where no relevant source exists. A system that handles unsupported questions safely is stronger than one that scores well only when the answer is present.

Plan freshness and deletion before launch

Choose how source changes reach the index: scheduled refresh, event-driven update, or another pipeline. When content is retired, remove or deactivate the corresponding indexed material so stale evidence does not remain available.

Version metadata helps when historical material must stay searchable. The model should not have to guess which document is authoritative.

Monitor production grounding quality

Track queries with poor results, missing sources, unexpected rankings, or answers that users correct. Those cases can reveal corpus gaps, chunking problems, stale content, or weak ranking.

Feed verified failures back into the evaluation suite. The retrieval system should improve from production evidence rather than relying only on a launch benchmark.

Keep query transformation visible

User questions do not always map directly to the best search query. A grounding pipeline may normalize terms, expand synonyms, extract filters, or issue several searches. Record those transformations so retrieval failures can be diagnosed instead of appearing as mysterious model behavior.

Query rewriting should be evaluated against known questions. More sophisticated rewriting is not automatically better if it drifts away from the user’s real intent.

Use source-specific chunking when content types differ

One chunking strategy rarely fits manuals, policy documents, tables, transcripts, and images equally well. Use structure-aware approaches where the source provides useful boundaries and preserve media-specific context when converting content into searchable representations.

The goal is consistent retrieval quality, not a single global chunk size.

Deduplicate equivalent content before ranking

Enterprise repositories often contain copied or syndicated material. If several near-identical versions occupy the top results, Claude receives less diverse evidence and may treat repetition as stronger support than it really is.

Use canonical-source rules, version metadata, or deduplication to keep retrieval results informative.

Use access filters early in the retrieval path

Do not retrieve restricted content and then ask the model to ignore it. Apply authorization before or during search so unauthorized evidence never enters the model context.

This reduces both security risk and unnecessary context use. It also makes audit simpler because the retrieved result set already reflects the caller’s permitted scope.

Measure citation correctness separately

A response can contain accurate claims but attach them to the wrong source. Test whether citations actually support the statement they accompany and whether the cited version is authoritative.

Citation evaluation is especially important in regulated or policy-heavy applications where source traceability is part of the product promise.

Use a retrieval fallback hierarchy

If vector search is degraded, the application may be able to use lexical search or a narrower repository rather than abandoning grounding completely. Define which fallbacks are acceptable and which situations should stop the response.

Fallback should be visible in telemetry so operators can see when users are receiving a reduced retrieval path.

Monitor unanswered questions as a corpus signal

Repeated questions with no supporting evidence may reveal a real knowledge gap rather than a search defect. Track them and route them to content owners. RAG can only ground answers in information the organization actually maintains.

This closes the loop between user demand and knowledge-base improvement.

Use different retrieval paths for different source classes

A product catalog, long policy document, image archive, and meeting transcript may benefit from different indexing and retrieval strategies. Keep the user-facing grounding contract consistent while allowing source-specific ingestion behind it.

This improves relevance without forcing every content type through one lowest-common-denominator pipeline.

Preserve a raw-source path for audit

Enriched or chunked representations are useful for search, but important workflows should still be able to trace a result to the original source. Store stable source identifiers and, where appropriate, offsets, page numbers, timestamps, or region metadata.

This makes citations, debugging, and source corrections much easier.

Use separate evaluation sets for retrieval and grounded response quality

A retrieval benchmark should answer whether the right evidence appears, while a grounded-response benchmark should answer whether the model used that evidence correctly. Keeping those sets separate makes it easier to diagnose whether a change in search, chunking, ranking, or prompting actually improved the system.

Use a smaller set of high-value end-to-end scenarios as a final integration check after both component tests pass.

A good RAG pipeline is an information system first

The model is only the final reasoning layer. Source governance, ingestion, extraction, indexing, retrieval, filtering, and evaluation determine which evidence the model sees.

If you can trace a document from source to chunk to index to retrieved evidence to grounded answer, you are reasoning at the level AI-103 expects.

  • img