TT Lab
Get started
Learn Learning paths Courses

Working With Customer Data

Building a Reusable Validator

Continue in TT Lab

Goal

You become able to build not a cleansing script that is written once and thrown away but a validator that automatically judges the files arriving every week.

Why it matters

The real test of a validator is not the broken files but the normal files. If the rules are so tight that even intact files fail, people start ignoring it every week with "it's probably that again", and two months later when a real problem comes, they ignore it in the same way. At that point having the validator is worse than not having one — because it makes people believe it is there while actually preventing nothing.

So a validator is always tested in two directions. Does it pass a normal file, and does it fail a broken file. And the result must be spoken through the exit code. 0 for pass, a nonzero value for failure. That is how it plugs straight into a batch pipeline.

There are four files in /opt/data/validate/. Three are broken in different ways and one is intact. The header of every file is id,customer,amount,region,date.

Three rules

Steps

  1. Create the /root/validate directory.
  2. Write the number of format-violating rows of bad_schema.csv in /root/validate/schema_bad.txt.
  3. Write the number of rows to be discarded as duplicates in bad_dupe.csv in /root/validate/dupe_rows.txt.
  4. Write the number of range-violating rows of bad_range.csv in /root/validate/range_bad.txt.
  5. Write the number of violating rows of good.csv in /root/validate/good_bad.txt.
  6. Write the sum of the four values in /root/validate/total_bad.txt.
  7. Create /root/validate/validate.sh. It takes the CSV path as its first argument, and must end with exit code 0 if there is no violation at all and with a nonzero code if there is even one.
  8. In /root/validate/report.md, write the names of the four files, the number of violations in each, and the total.

Notes

Create the working directory

Create the /root/validate directory.

Gather the results under /root/validate.

Count the format-violating rows

Write the number of format-violating rows of bad_schema.csv in /root/validate/schema_bad.txt.

It is the number of rows in bad_schema.csv that break at least one of three: the field count, an empty customer name, and a non-integer amount.

Count the duplicate rows

Write the number of rows to be discarded as duplicates in bad_dupe.csv in /root/validate/dupe_rows.txt.

It is the number of rows in bad_dupe.csv where an id that appeared before appears again. The row that appeared first is not counted.

Count the range-violating rows

Write the number of range-violating rows of bad_range.csv in /root/validate/range_bad.txt.

These are the rows in bad_range.csv where amount falls outside 1 to 10000000. 0 is also a violation.

Check the normal file

Write the number of violating rows of good.csv in /root/validate/good_bad.txt.

good.csv must have no violations. If you do not get 0 here, the validator is over-detecting.

Compute the total of violations

Write the sum of the four values in /root/validate/total_bad.txt.

Add up all the numbers that came out of the four files.

Build the validation script

Create /root/validate/validate.sh. It takes the CSV path as its first argument, and must end with exit code 0 if there is no violation at all and with a nonzero code if there is even one.

It takes the file path as an argument, and must end with exit code 0 if it passes and with a nonzero value if there is a violation.

Write the validation report

In /root/validate/report.md, write the names of the four files, the number of violations in each, and the total.

It must contain the four file names, the number of violations in each, and the total.