WORKED EXAMPLE · V1.0.0
Find the duplicate.
Keep the evidence.
This is the package’s synthetic CSV fixture, not a customer dataset or an automatic repair result. The source contains four data records.
Choose explicit rules
{"required":["sku","quantity","unit_cost","updated_at"],
"unique":["sku"],
"types":{"sku":"text","quantity":"integer","unit_cost":"decimal","updated_at":"date"}}The text ID 0012 keeps its leading zeros. Required cells must be non-empty; unique keys are exact and case-sensitive.
Run the checker
python3 scripts/audit_csv.py examples/broken.csv --rules examples/rules.json --json broken.json --output broken.mdExit code 1 · 4 records · 6 findings:
- Record 2 / physical line 3: duplicate sku, first seen in record 1; invalid integer quantity, decimal unit_cost and date updated_at.
- Record 3 / physical line 4: missing required sku cell.
- Record 4 / physical line 5: only 3 columns where 5 are expected.
JSON records column names and row locations; reports contain no cell values or input paths. Schema and duplicate locations can still be sensitive.
Confirm a correction, then rerun
The included corrected.csv is a manually prepared illustration with a distinct key, valid types and complete rows. Its four records pass with exit 0 and zero findings. These replacements are invented for this example; confirm real replacement values with the data owner.
python3 scripts/audit_csv.py examples/corrected.csv --rules examples/rules.json --json corrected.json --output corrected.mdThe original file is preserved. Compare input hashes, row counts and findings before accepting your own corrected copy. Reports require new destinations.
Remaining checks
Check your importer’s conventions and business constraints separately. A pass does not prove inventory quantities, accounting values or target-database compatibility. No uploads, repairs or import execution occur.
Get CSV Data Quality Audit · $4.90 ↗