Our eval set is 3,000 documents and 14 model calls per document. That’s 42,000 calls per candidate.
Where the money went
About a third of our spend was re-running calls that had already succeeded, because one grader timed out near the end of a batch.
Memoize everything
Once each call became a step, retries and replays read finished results back for free. Our spend per candidate fell 38%.



