Systems and data
Model evaluation writer
Designing the tests that show whether an AI system is good enough for a specific real use, and writing down honestly where it fails.
Written by Nivaan, founder of OffLadder · Last reviewed
Test it this week
Pick a task you know well and write ten test prompts for it, including three you expect to break a model. Score the answers against your own rubric, then write a paragraph a colleague could act on.
What the work actually contains
- Collecting the awkward cases a system will meet in the wild, not the tidy ones
- Writing what a good answer looks like for each, in enough detail that two people would score it the same way
- Running the tests as the system changes and reporting the regressions nobody wants to hear about
- Turning the findings into plain language for people who will not read a chart
Why it is appearing now
Anyone deploying a model has to answer 'how do you know it works here?', and generic benchmarks do not answer it for a specific job.
What AI changes about it
The subject matter is AI, but the craft is old: careful test design and honest reporting. Models can generate test cases; deciding what counts as correct still needs a person who knows the domain.
The human abilities it leans on
- Precision about what 'correct' means in a specific context
- Comfort delivering unwelcome findings clearly
- Curiosity about edge cases and the willingness to hunt for them
What it can grow out of
- Teaching and marking, where you already write rubrics
- QA, editing, auditing, compliance
- Any domain expertise — the scarce part is knowing the field, not the tooling
What nobody knows yet
How much of this becomes automated scoring and how much stays with domain experts who define what good means.
Guides worth reading