AI · document intelligence

An eval pipeline that makes 14 model calls per document now pays for each call once, even when a grader flakes.

38% lower eval spend

14 model calls per document

6x more eval runs per week

Memoized steps cut our eval bill by 38%.

Hana Sato

ML engineer

kestrel.ai

kestrel.ai case study cover

Where they started

kestrel.ai extracts structured data from legal documents. Every prompt change runs against 3,000 labelled documents, and each document takes 14 model calls to extract and grade.

What they built

The eval runs as one parent workflow that invokes a child per document. Each generation and grading call is a step, so when step 12 times out the retry reuses the 11 results already saved.

Provider rate limits are handled with a concurrency key per provider and model, which let them run two providers side by side without tripping either.

What changed

Eval spend per candidate fell 38%, which made it cheap enough to evaluate every prompt change instead of batching them weekly.

Get started

Ship the agent. Keep the receipts.

Free for 50,000 steps a month. No credit card, no separate workers, and your first durable workflow deployed before lunch.

orrindel

Durable runtime for AI agents and automations. Every run, in plain text.

Book a 20-minute demo →
All systems normal
99.99% uptime · last 90 days
90 days agotoday
regions us-east · eu-west · ap-south
soc 2 type ii · gdpr · hipaa (baa)
sdk v4.2.1 · node · python · go
© 2026 Orrindel Labs, Inc.PrivacyTermsSecurityMade in plain text.

Create a free website with Framer, the website builder loved by startups, designers and agencies.