The problem
An eval run makes thousands of model calls. When one grader times out at 80% complete, re-running from scratch burns money and hours.
How the workflow runs
A new candidate triggers the run, which fans out over the dataset with a concurrency limit per provider. Each generation and each grading call is a memoized step, so a retry or replay reads finished results back for free.
Scores stream into a comparison table as they land; when the last batch finishes, the run posts a summary with regressions highlighted.
What changes
Evals become cheap enough to run on every prompt change, and flaky graders stop quietly skewing your numbers.



