Systems and data
Synthetic data curator
Making and checking the artificial examples used to train and test systems, and being the person who notices what the data quietly leaves out.
Written by Nivaan, founder of OffLadder · Last reviewed
Test it this week
Generate 50 example records for something you understand well, then list every way in which your 50 are unrealistically similar. Stop when you have five.
What the work actually contains
- Deciding what a dataset needs to represent, including the rare cases
- Generating examples, then hunting for the ways they are all subtly alike
- Documenting what the data can and cannot be used to claim
- Working with the people affected by the system to find the cases they would want included
Why it is appearing now
Real data is often restricted, expensive or thin in exactly the cases that matter, so synthetic examples are being used to fill the gap — which creates a new way to be wrong.
What AI changes about it
Generation is cheap now; judgement about representativeness is not. The value has moved to the curating.
The human abilities it leans on
- Statistical scepticism without needing to be a statistician
- Willingness to look for the missing case rather than the confirming one
- Clear documentation habits
What it can grow out of
- Research assistance, survey work, data entry with an eye for anomalies
- Librarianship and archives
- Any field where you know what the unusual cases look like
What nobody knows yet
How much of this becomes tooling and how much stays judgement.
Guides worth reading