Data cleaning & validation pipeline
Attach a raw dataset and the agent profiles, cleans, and validates it against explicit rules, then drafts the cleaned data and a quality report for you to approve.
Saves you ~4.3 h / run
How it works
- Trigger
- When you start the “Clean and validate data” workflow.
- Job
- Profile, clean, validate, and quality-check the dataset.
- Outcome
- A cleaned, validated dataset with exception records and quality evidence.
What it installs
Agents 2
-
Data Engineer Agent
Cleans and validates data, QAs its own output, and drafts a quality report.
-
Data cleaning & validation pipeline Quality Reviewer
Checks primary evidence, domain controls, deliverable completeness, and communication quality, stopping the run when the work is wrong, unsupported, incomplete, or uncertain.
Teams 1
-
Data cleaning & validation pipeline quality team
The delivery agents produce the work while an independent quality reviewer checks each workflow handoff against explicit evidence, domain, and communication requirements before the run can continue.
Workflows 1
-
Clean and validate data
Profile, clean, validate, QA the output, then draft the cleaned data and quality report for approval.
Documents 1
-
Validation rules
Your editable config of the checks, tolerances, and exclusions that define when a dataset counts as clean. Set these to steer the pipeline.
Goals 1
-
Only clean data flows downstream
Only validated, reproducibly cleaned data enters downstream analysis. Success looks like: Every dataset closes with an approved cleaned dataset and a quality report that records each validation check, its result, and a confidence rating.
Skills 3
-
validate-data
QA a dataset and the analysis built on it before it is trusted — methodology, calculation, and bias checks producing a confidence assessment. Adapted from anthropics/knowledge-work-plugins/validate-data.
-
data-context-extractor
Profile messy data and extract its structure, entities, metrics, and hidden hygiene rules before transforming it. Adapted from anthropics/knowledge-work-plugins/data-context-extractor.
-
etl-pipeline
Design a repeatable extract-transform-load flow with cleaning, validation, idempotency, and quality reporting. Adapted from claude-office-skills/skills/etl-pipeline.
Folders 1
-
Data Cleaning
Requirements
- Product analytics access — Connect your product analytics so the agent can profile and cross-check source tables directly. Until you connect it, it works from attached exports and the validation rules document.
Setup guide
How to Automate a Data Cleaning Pipeline
A practical guide to cleaning datasets with repeatable transforms, validation rules, QA, and human approval.
Read the setup guideDon't see your workflow? Describe it.
A sentence or two about a recurring job is enough. We design the playbook that runs it and show you exactly what it saves.