RAG Pipelines for Databricks GenAI Engineer
RAG is one of the clearest through-lines in the Databricks Generative AI Engineer Associate exam. The current guide expects candidates to select source documents, extract content, remove low-value material, choose chunking strategies, write prepared data to Delta tables in Unity Catalog, evaluate retrieval, use reranking, assemble a RAG application, and monitor performance after deployment.
The general embeddings and RAG guide explains the architecture. This page keeps the focus on the Databricks pipeline and the decisions that connect data preparation to Vector Search, agents, MLflow, and production governance.
Source selection should follow the user questions and business scope. A support application may need product manuals, policies, release notes, and account-specific structured data but not every document the organization owns.
More content can reduce quality when irrelevant or conflicting material crowds the retrieval stage.
Identify current versions, owners, licensing, permissions, and update frequency. Historical or duplicated documents should be tagged so the application can distinguish them from authoritative content.
RAG quality depends on source governance before any embedding is created.
HTML, PDF, scanned images, office documents, and structured files can require different extraction packages or services.
The current exam guide explicitly tests choosing an appropriate extraction approach from the source format. Extraction quality should be inspected before chunking.
Navigation, repeated headers, boilerplate, legal footers, irrelevant sidebars, and duplicated sections can dilute semantic signals.
Filtering should preserve evidence that users need while removing text that consistently produces poor retrieval.
Headings, paragraphs, tables, semantic sections, and token constraints all influence chunking.
Large chunks preserve context but reduce granularity. Small chunks improve precision but can fragment meaning and create many more records.
Overlap can preserve context across chunk boundaries, but excessive overlap increases index size and duplicates evidence.
Choose overlap based on the content structure and measured retrieval behavior rather than a default percentage.
Store text, source identifier, chunk identifier, metadata, timestamps, and other fields needed for filtering and lineage.
Unity Catalog provides the governance context so retrieval assets inherit clear ownership and access rules.
Customer, region, product, language, date, document status, or access group can all be important filters.
Apply filters before the model receives content so unauthorized or irrelevant evidence does not enter the context.
Create questions with known relevant documents or chunks and measure whether the retriever finds them.
Without labeled queries, chunking and index changes become subjective and teams can mistake plausible results for systematic improvement.
Recall, ranking position, precision-style measures, and task-specific judgments can show whether the right evidence is present and prominent.
Keep retrieval evaluation separate from final-answer evaluation so you know which layer to improve.
Structure-aware, semantic, parent-child, or other advanced chunking patterns can improve quality for some corpora.
Use them when ordinary chunking fails in a measurable way; additional complexity is not valuable by itself.
A first-stage retriever can produce a broader candidate set and a reranker can apply a stronger relevance judgment before context assembly.
Reranking cannot recover a document that was never retrieved, so evaluate recall first.
Pass a focused set of evidence into the model and retain source metadata so the response can be traced or cited.
Do not collapse several sources into an anonymous text block if source authority matters.
Account state, transaction status, inventory, or other fast-changing facts may be better retrieved through structured tools than embedded documents.
A robust RAG application can combine document evidence with live data without confusing the two authority sources.
A tool-using agent may decide when search is needed and can combine retrieval with other actions.
Keep retrieval permissions scoped to the user’s or application’s allowed data.
Check whether the answer is supported by the retrieved evidence, whether important claims are complete, and whether the application handles missing evidence safely.
The evaluation framework helps separate retrieval quality from generation quality and business task success.
Production RAG can degrade when ingestion jobs fail, source documents change, embeddings become stale, or filters no longer match the business model.
Track the age and completeness of indexed content alongside user-facing quality.
Check document ownership, freshness, duplication, formatting, encoding, language, and whether the content actually contains the answers users need.
Indexing a low-quality corpus faster only produces low-quality retrieval faster.
Prepared chunks stored in Delta allow engineers to inspect text, metadata, source IDs, and transformations before search hides those details behind an index.
This makes debugging and reprocessing easier when extraction or chunking changes.
Build two prepared versions of the same corpus and run the same labeled queries against them.
Compare recall, ranking, vector count, and latency so chunking becomes an evidence-based choice.
Reranking adds latency and sometimes cost. Measure how often it moves the correct evidence into a useful position.
If first-stage retrieval already performs well, the extra component may not be justified.
Unity Catalog and application filters should ensure users or agents only retrieve evidence they are permitted to access.
Do not treat RAG as a new path around existing data governance.
Repeated zero-result or poor-result queries can reveal missing sources, vocabulary gaps, or wrong metadata.
Use privacy-safe analytics to improve the corpus and retrieval rules over time.
A new embedding model, chunking strategy, or index configuration can change answer quality broadly.
Validate before cutover and keep the prior retrieval configuration available until the new one proves stable.
Metadata such as heading, document type, product, version, and owner can improve filtering and interpretation even when the semantic text is embedded.
Do not throw away source structure simply because Vector Search can find similar text.
Repeated copies of the same policy or manual can dominate retrieval and make frequency look like authority.
Deduplicate or identify canonical versions during preparation so results remain diverse and trustworthy.
Carry document and chunk identity into the final context so the application can cite or audit the evidence used.
This supports debugging and helps subject-matter experts judge whether the response came from the right source.
More chunks affect embedding jobs, index size, refresh time, query latency, and potentially reranking cost.
Capacity planning should consider the expected corpus after growth, not only the initial pilot.
Document policy questions may use Vector Search, while current account or inventory facts may come from structured tools.
Routing questions to the right evidence source improves both freshness and explainability.
Document extraction, chunking, embedding, and enrichment do not need to occur inside the user request. Run them as controlled background pipelines and expose only the prepared index to the interactive application.
This keeps user latency predictable and makes ingestion failures easier to monitor and retry.
Some questions need evidence from more than one chunk or document. Include those in the evaluation set so retrieval is not optimized only for single-passage answers.
Check whether the context assembler preserves source boundaries and whether the model combines evidence without inventing unsupported links.
If two current sources disagree, preserve the conflict rather than allowing the model to average them into one confident statement.
Source authority, date, and ownership metadata can help the application decide whether one source should dominate or whether the user should see the disagreement.
Record which source version, extraction method, chunking configuration, and embedding model produced each indexed representation.
When quality changes or a source is corrected, lineage lets the team identify which chunks need reprocessing instead of rebuilding blindly.
The value of retrieval is selecting the evidence needed for one request rather than copying the knowledge base into every model call. A focused context is easier to evaluate, cheaper to process, and less likely to contain conflicting material.
Queries with poor results can reveal missing sources, weak chunking, terminology gaps, or ranking problems.
Add privacy-safe versions of those cases to the evaluation set so the pipeline improves from real usage.
