Databricks Data Engineer Professional: Data Quality Engineering

Data quality on Databricks is an engineered contract between producers, pipelines, and consumers. The Professional exam expects more than knowing how to filter nulls. Candidates need to reason about schema, validity, uniqueness, freshness, referential integrity, CDC behavior, expectations, quarantine, observability, and the operational decision that follows a failed rule.

The Databricks Certified Data Engineer Professional scope places quality inside production pipeline design. A high-quality pipeline does not just detect bad data; it decides whether to warn, drop, quarantine, fail, or repair it in a way that preserves downstream trust.

The most useful mindset is to define quality at each layer. Raw capture has different obligations from a curated analytical table, and a rule that belongs in silver can be harmful if it silently deletes evidence in bronze.

Define quality dimensions before writing checks

Quality rules should correspond to real consumer risk. Completeness asks whether required values are present. Validity checks domains, formats, and ranges. Uniqueness enforces identifiers or business keys. Consistency compares related fields or systems. Timeliness measures freshness and lateness. Referential integrity checks whether relationships point to valid entities.

A generic “not null on every column” policy is rarely useful. Some fields are legitimately unknown, while one missing business key may make the record unusable. Start with the downstream decision and identify which defects can change that decision.

The data engineer skill map is a good reminder that quality sits beside modeling, pipelines, governance, and operations rather than being a separate cleanup phase.

Schema is the first quality contract

Schema enforcement protects downstream code from incompatible structure. Types, required fields, and naming conventions should be explicit where stability matters. Schema evolution can accept compatible additions, but it should not turn every upstream change into an automatic production change.

Distinguish structural compatibility from semantic compatibility. A field can keep the same string type while changing meaning, allowed values, or units. No schema engine can detect every semantic break. That requires producer contracts, metadata, tests, and monitoring.

In bronze, preserving source fidelity may justify a permissive capture pattern. In silver and gold, stronger contracts are usually appropriate because consumers assume more trust.

Lakeflow expectations make data-quality behavior declarative

Lakeflow pipeline expectations attach a boolean condition to a dataset and specify what happens when records violate it. The response can warn while retaining the record, drop violating records, or fail the update. That makes quality policy visible next to the transformation instead of hiding it in scattered conditional code.

Choose the action based on business consequence. A malformed optional attribute may warrant a warning. A record with an impossible primary key may need quarantine or rejection. A financial aggregation built on invalid currency or date values may justify failing the update to protect downstream reporting.

The event log can expose expectation metrics, which makes trends visible over time rather than reducing quality to a one-time pass/fail result.

Quarantine preserves evidence while protecting consumers

Dropping a bad record removes it from the curated output, but it can also erase the evidence needed for investigation. A quarantine pattern routes invalid records to a controlled table with failure reasons, source metadata, and ingestion context so operators can diagnose and, where appropriate, repair or replay them.

Quarantine is especially useful when upstream correction is possible. Instead of manually editing the curated table, fix the producer or transformation, then reprocess the quarantined input through the same rules. This keeps the lineage and audit trail intact.

Do not let quarantine become a permanent landfill. Track backlog size, age, ownership, and recurring failure categories so the quality system drives source improvement.

Deduplication must represent business identity

Duplicate files, retry events, CDC records, and repeated source extracts can all produce multiple physical rows for one logical fact. Deduplication therefore requires a stable key and ordering rule. In streaming workloads, watermarks can bound how long duplicate keys are retained in state.

For CDC, the problem is more nuanced. Multiple versions of the same entity may be legitimate. Sequence information determines which update is current, and late-arriving changes may need special handling. AUTO CDC can simplify common SCD patterns, but the engineer still needs correct keys and sequencing semantics.

A deduplication rule that removes valid history is a data-quality defect, even if it reduces row count.

Medallion architecture can express progressive trust

A useful bronze layer captures the source with enough metadata to support replay and investigation. Silver applies normalization, type cleanup, deduplication, key validation, and cross-source reconciliation. Gold applies business-level calculations and exposes data at a consumer-ready grain.

This progression makes it possible to trace a bad gold metric back to the curated inputs and then to source records. Medallion architecture and quality progression is valuable precisely because each layer has a different trust promise.

Avoid placing irreversible cleaning in the raw layer. If a rule later proves wrong, you want the original evidence available for reprocessing.

Quality metrics need thresholds, trends, and ownership

A quality check becomes operational only when someone knows what to do with the result. Define thresholds for null rate, duplicate rate, freshness, invalid domain values, reconciliation differences, or other critical metrics. A single bad record in ten million may be acceptable for one dataset and catastrophic for another.

Trend the metrics. A slow rise in rejected records can reveal upstream drift before a hard failure occurs. Compare by source, partition, customer, or other meaningful segment so localized problems do not disappear inside a global percentage.

Every critical rule needs an owner and an escalation path. Data quality without ownership is just logging.

Testing and monitoring should cover both code and data

Unit tests can verify transformation logic on controlled examples. Integration tests prove that components work together. Data tests verify assumptions about production inputs and outputs. Monitoring then checks those assumptions continuously after deployment.

The broader pipeline architecture should expose where each test belongs. A source freshness check belongs near ingestion; a business aggregation reconciliation belongs near the curated output.

When a test fails, preserve the failing input and expected contract so the incident becomes reproducible rather than anecdotal.

Imagine a customer master dataset where 0.2 percent of records lack country codes, a small source suddenly emits duplicate customer IDs, and one upstream application begins sending timestamps in local time instead of UTC. A simple pass/fail rule cannot treat all three defects the same way. Missing optional geography may warrant warning, duplicate primary identities may require quarantine, and the timestamp contract change may justify failing the update because it corrupts temporal logic downstream.

This scenario shows why quality severity belongs to business impact. Define which rules protect identity, money, regulatory reporting, and other critical outcomes, then choose warn/drop/fail behavior accordingly. Less critical defects can remain visible without stopping the entire pipeline. Critical defects should stop bad data before it crosses a trust boundary where many consumers will assume it is valid.

Quality recovery should be reproducible. If a source team fixes the customer-ID duplication, replay the quarantined or affected source range through the same transformation and expectations. Avoid manually editing the silver table to make metrics look clean. Manual correction breaks lineage and makes it difficult to prove which version of the business rule produced the current data.

Treat quality metrics as a product. Publish freshness, valid-row rate, duplicate rate, expectation failures, and unresolved quarantine backlog where appropriate. A consumer should be able to tell whether data is merely available or actually within its trust objective. Professional data engineering turns quality from an invisible cleaning step into an observable service contract.

Quality rules should be versioned with the transformation logic that depends on them. If a domain rule changes from allowing ten status values to allowing twelve, the code, expectation definition, documentation, and test data should move together. This makes a quality change reviewable and reproducible. It also gives operators a way to explain why yesterday’s data passed and today’s data failed even when the source did not change. Unversioned rules turn quality incidents into archaeology.

Reconciliation catches a different class of problem from row-level validation. A pipeline can produce individually valid records while losing 8% of transactions because one source partition was skipped. Compare source and target counts, sums, distinct keys, or control totals at meaningful boundaries. Reconciliation is especially important after joins, aggregations, CDC merges, and backfills because those operations can be logically wrong without triggering schema or null checks. Build tolerances where exact equality is not meaningful, but make the tolerance explicit.

Data quality also has a timeliness dimension. A perfectly valid table that is six hours late may be unusable for an operational dashboard. Track source arrival, ingestion completion, transformation completion, and published freshness separately so the team can locate the delay. Freshness SLAs should name the business owner and escalation path just like availability SLOs. When late data is acceptable, distinguish ‘late but complete’ from ‘current but incomplete’ because consumers may prefer different behaviors.

For critical datasets, quality incidents need the same discipline as software incidents. Record affected tables and consumers, the violated rule, the first bad partition or event, the containment action, the correction method, and the prevention change. If bad data was already consumed, recovery may require downstream recomputation or communication rather than only fixing the upstream table. This is why lineage and ownership are operational quality controls, not catalog decoration.

The Professional exam perspective is therefore broader than writing an expectation expression. You should be able to decide where the rule belongs, what should happen when it fails, which evidence proves the impact, and how the system returns to a trusted state. That combination of policy, measurement, and recovery is what turns validation into data quality engineering.

Quality engineering should also distinguish producer defects from consumer-specific expectations. A source may be valid according to its contract while still being unsuitable for one downstream model. Keep universal source rules near ingestion, and place business-specific checks closer to the curated product that needs them. This prevents one consumer’s assumptions from blocking unrelated workloads while still making critical expectations enforceable. The placement of a rule is therefore part of its design: it determines who owns the failure, who is affected, and which layer must be corrected.

Professional quality engineering minimizes silent failure

The most dangerous data-quality problem is not always a failed job. It is a successful job that produces believable but wrong data. That is why schema checks, expectations, reconciliations, freshness metrics, and domain validation need to cover conditions that the execution engine cannot detect on its own.

A strong design also makes exceptions visible. If a team intentionally overrides a rule, record who approved it, why, how long the exception lasts, and which consumers may be affected. Temporary quality exceptions that become permanent create hidden debt.

For the Professional exam, think like an owner of a data product: define what trusted means, make violations observable, preserve evidence, and recover through controlled reprocessing rather than manual correction.

  • img