Computing Library › Data Systems
Data Systems

Data-Quality Validation

Data-quality validation checks data against explicit expectations for completeness, validity, and consistency, catching problems before they reach consumers.

Quality as measurable expectations

Data quality is not a vague virtue; it is a set of checkable claims about data. Validation makes those claims explicit as expectations and tests every batch against them, so a broken upstream feed or a logic bug is caught at the boundary rather than discovered later in a wrong result. The alternative, trusting data silently, means quality problems surface as inexplicable downstream errors.

Dimensions of quality

Kronos motion — validation

Where checks run

Validation runs at ingestion (reject or quarantine bad input at the door), between pipeline stages (catch a transformation that introduced errors), and before publication (a final gate on what consumers see). Placing checks at these boundaries localizes failures: you learn which stage broke, not merely that the final output is wrong.

Fail loud or quarantine

When a check fails, the pipeline must decide. Fail-fast stops the run and alerts, appropriate when bad data must never propagate. Quarantine diverts failing records to a side location and processes the rest, appropriate when partial output is better than none. The right choice depends on whether consumers prefer correctness or availability, and should be a deliberate policy rather than an accident of implementation.

Distribution checks and drift

Beyond per-record rules, statistical checks compare a batch's distribution to a historical baseline: a sudden shift in a column's mean or null rate signals a problem even when every individual value is technically valid. This is the same drift detection used to monitor model inputs. Expectations should be versioned alongside the pipeline so a schema or logic change updates them deliberately. See expectation suites and data contracts.