Where they started
kestrel.ai extracts structured data from legal documents. Every prompt change runs against 3,000 labelled documents, and each document takes 14 model calls to extract and grade.
What they built
The eval runs as one parent workflow that invokes a child per document. Each generation and grading call is a step, so when step 12 times out the retry reuses the 11 results already saved.
Provider rate limits are handled with a concurrency key per provider and model, which let them run two providers side by side without tripping either.
What changed
Eval spend per candidate fell 38%, which made it cheap enough to evaluate every prompt change instead of batching them weekly.


