Durable runtime for AI agents · v4.2

AI workflows thatshow their work.

Orrindel runs your agents and automations as durable steps. Every call is retried, every decision is logged, and any run can be replayed from the step that broke.

Start building →
Steps run today
48,123,379
Retries absorbed this week
212,904
p99 step overhead
0 ms
Runs lost to crashes
0
Runs in production at
◆Halvard Health▲NORTHPIER●kestrel.ai■Tallow & Co◐Quanta Freight✦meridel▮BRIGHTWELL◇sable/ops◆Oxbow Legal▲Pellam
01 · The problem

Most agent code forgets everything when it crashes.

A script holds its progress in memory. When the connection drops on step five, the restart runs steps one to four again, and your customer gets refunded twice.

Steps redone
0
Duplicate refunds
0
Time lost
0s
refund-agent.js · pm2 logsidle
09:14:02.118$ node refund-agent.js --ticket 48213
09:14:02.402[1/6] classify ............ ok intent=refund (1.21s)
09:14:03.611[2/6] fetch account ....... ok ac_7Q2 (0.33s)
09:14:03.944[3/6] find charges ........ ok ch_31x duplicate (0.54s)
09:14:04.487[4/6] issue refund ........ ok re_88k $49.00 (0.27s)
09:14:04.760[5/6] draft reply (llm) ... …
09:14:41.003Error: read ECONNRESET at TLSSocket.onStreamRead
09:14:41.004process exited with code 1
09:14:46.220pm2: restarting refund-agent (attempt 2)
09:14:46.531[1/6] classify ............ ok (again · 1.19s · $0.0009)
09:14:47.722[2/6] fetch account ....... ok (again)
09:14:48.061[3/6] find charges ........ ok (again)
09:14:48.603[4/6] issue refund ........ ok re_91c $49.00 ← second refund
# the customer now has two refunds and one confused support agent.
02 · Durable execution

When a step breaks, the run doesn’t start over.

Orrindel saves the result of every step. Fix the bug, press replay, and the run picks up at the step that failed, with everything before it intact.

  1. 01RunFive steps start. Each result is saved as it lands.
  2. 02Failissue-refund sends no amount. Retries can’t fix a bug, so it stops at 3/3.
  3. 03FixChange one token. Deploy. The failed run is still waiting.
  4. 04ReplaySteps 1–3 come back from memo in 0 ms. Only the broken step runs again.
Steps restored
0
Time not re-spent
0.00s
LLM calls repeated
0
Duplicate refunds
0
refund-flow.tsrun_01JAX9TQticket #48213● running
1export const refundFlow = orrindel.workflow(
2 { id: "refund-flow", trigger: "ticket.created" },
3 async ({ event, step }) => {
4
5 const intent = await step.run("classify", () => llm.classify(event.body))
6 const account = await step.run("fetch-account", () => crm.get(event.customerId))
7 const charges = await step.run("find-charges", () => billing.list(account.id))
8 const refund = await step.run("issue-refund", () =>
9 billing.refund(charges.duplicate.id, charges.duplicate.amount))
10 await step.run("reply", () => mail.send(event.from, refunded(refund)))
11 }
12)
StepTimeSaved output
●classify0.00s
·fetch-accountqueued
·find-chargesqueued
·issue-refundqueued
·replyqueued
attempt log
replay
—
0s1s2s3s4s
03 · Primitives

Six calls cover almost every workflow you’ll write.

No new language and no YAML. You write normal functions and wrap the parts that talk to the outside world.

step.run()01

Run it once, keep the result

Wrap any call. The result is saved, so a retry or replay reads it back instead of calling again.

const acct = await step.run("fetch-account",
  () => crm.get(event.customerId))
step.sleep()02

Wait days without a server

Pause a run for minutes or months. Nothing polls and nothing is billed while it sleeps.

await step.sleep("trial-ends", "14d")
await step.run("send-offer", sendOffer)
step.waitForEvent()03

Pause until something happens

Hold a run until a matching event arrives, like a paid invoice or a signed form, with a timeout.

const paid = await step.waitForEvent("paid",
  { event: "invoice.paid", timeout: "7d" })
retries04

Retry with backoff, by default

Failed steps retry on their own: 3 times, with exponential backoff. Errors you mark final stop at once.

step.run("charge", charge, {
  retries: 3, backoff: "exponential" })
concurrency05

Limit how much runs at once

Cap parallel runs per key, like one per customer, so a burst from one account never starves the rest.

orrindel.workflow({ id: "sync",
  concurrency: { limit: 2, key: "event.accountId" } })
step.invoke()06

Call another workflow

Compose workflows. The parent waits for the child's result, and both show up in one trace.

const lead = await step.invoke("enrich",
  { function: enrichLead, data: { email } })
04 · Human in the loop

Let a person say yes before the agent does anything permanent.

Any step can pause for a decision. The run waits, for seconds or for days, then carries on down the branch the person picked. Every answer is recorded with who gave it and when.

Timeout
24h, then escalates to #support-leads
Channels
Slack, Teams, email or your own UI
Audit
who, when, and what they saw
Cost while waiting
$0.00
orrindel#support-approvals09:22

Refund needs approval: $49.00 to Jamie Alvarez for ticket #48213. The March charge appears twice.

amount
$49.00 USD
charges
ch_31x · ch_30q (Mar 3)
agent confidence
0.94
policy
refund window 60d ✓
◷ expires in 23:59:59
run_01JAX9TQrefund-flow
✓draft_reply1.60s · 182 words
◷wait_for_approvalwaiting · 0s
·send_refundqueued
·close_ticketqueued
05 · Flow control

One noisy customer shouldn’t slow down everyone else.

Set a concurrency limit and a fairness key, and Orrindel serves each account in turn. A burst from your biggest customer no longer puts the smallest one at the back of the line.

acct_bigco · 40 jobs
0/40 done · waited p95 0.0s
acct_ridge · 8 jobs
0/8 done · waited p95 0.0s
acct_tern · 3 jobs
0/3 done · waited p95 0.0s
t = 0.0srunning 1/4 queued running doneconcurrency: { limit: 4, key: "event.accountId" }
06 · Observability

See every prompt, tool call and token, in the order it happened.

Each run keeps a full trace: what the model was asked, what it answered, which tools it called and what that cost. Hover a span to open it.

span0s2s4s6sms
support-triage6800
└ classify1200
└ search_docs800
└ embed_query170
└ fetch_account420
└ draft_reply3200
└ lookup_export340
└ requeue_export220
└ grade_reply700
└ send_reply500
LLM calltool calleval
draft_replyLLM · 3.20s
model
large
tokens
2,904 → 312
cost
$0.0031
tool calls
2
INPUT
System: You are a support engineer… Context: 4 passages, account ac_2Hv (business, eu-west)…
OUTPUT
"Hi Priya, large exports over 2 GB were timing out in eu-west this morning. We’ve raised the limit and re-queued yours…"
spend · last 24h$50.42 total1.9M tokenspeak 14:00 · $4.60budget alert at $80/day · ok
00:0006:0012:0018:0023:00
07 · Tools

Every integration is a typed function your agent can call.

Connect CRMs, billing, mail, chat, databases and browsers as typed tools. The types become the schema the model sees, so a bad tool call fails before it runs, not after.

crm· cached 60s
crm.get(customerId: string)
→ Promise<Account>
billing· idempotent
billing.refund(chargeId: string, amount: Money)
→ Promise<Refund>
mail· retries 3
mail.send(to: Email, body: Markdown)
→ Promise<MessageId>
chat· rate 1/s
chat.post(channel: string, text: string)
→ Promise<Message>
tickets· idempotent
tickets.update(id: number, patch: TicketPatch)
→ Promise<Ticket>
vector· p95 180ms
vector.search(query: string, k?: number)
→ Promise<Passage[]>
sheets· retries 5
sheets.append(sheetId: string, row: Cell[])
→ Promise<RowRef>
calendar◷ needs approval
calendar.book(slot: TimeRange, who: Email[])
→ Promise<Event>
payments· idempotent
payments.capture(intentId: string)
→ Promise<Payment>
docs· retries 3
docs.create(title: string, body: Markdown)
→ Promise<DocRef>
sql· read-only role
sql.query(text: SQL, params: unknown[])
→ Promise<Row[]>
browser· timeout 90s
browser.open(url: URL, actions: Action[])
→ Promise<PageSnapshot>
pdf· retries 2
pdf.extract(file: FileRef, schema: Zod)
→ Promise<T>
http· circuit breaker
http.request(req: RequestInit & { url })
→ Promise<Response>
storage· idempotent
storage.put(key: string, data: Blob)
→ Promise<ObjectRef>
sms◷ needs approval
sms.send(to: Phone, text: string)
→ Promise<SmsId>
queue· exactly once
queue.publish(topic: string, msg: Json)
→ Promise<Offset>
llm· tokens metered
llm.complete(prompt: Prompt, tools?: Tool[])
→ Promise<Completion>
embeddings· batched 96
embeddings.create(input: string[])
→ Promise<Vector[]>
webhooks· deduped 24h
webhooks.receive(source: string, verify: Secret)
→ Promise<Event>
120+ ready-made tools·bring your own with one function·MCP servers supported
08 · Your language

Write it in the language you already ship.

Workflows live in your repo, deploy with your app and run on your cloud or ours. TypeScript, Python and Go SDKs share one runtime and one dashboard.

workflows/onboarding.ts
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
  1. STEP 01
    Install the SDK
    $ npm i @orrindel/sdk
    or pip install orrindel · go get orrindel.dev/sdk
  2. STEP 02
    Run the dev server
    $ npx orrindel dev
    Local dashboard on :8288, with replay and traces
  3. STEP 03
    Deploy with your app
    $ git push
    Workflows register on boot. No separate workers.
10 · From the teams

What engineers wrote in the commit that switched them over.

$ git log --author-date-order --grep=orrindel
a41f9c2chore: delete our home-grown queue and retry layerMaya Okafor · 3 days ago

“We moved claims intake onto Orrindel in two sprints. The pager for stuck jobs used to go off weekly. It hasn’t gone off since March.”

Maya Okafor · Staff engineer, Halvard Health
+212 −2,318
++----------------------

11 · Security

Built for the teams your security review will send.

Your data stays in the region you pick, encrypted with keys per workspace. Enterprise can run the whole runtime inside your own cloud account.

Compliance

Compliance

SOC 2 Type II, GDPR, HIPAA with a signed BAA

SOC 2 Type II, GDPR, HIPAA with a signed BAA

Data residency

Data residency

us-east, eu-west or ap-south, pinned per workspace

us-east, eu-west or ap-south, pinned per workspace

Encryption

Encryption

AES-256 at rest with per-workspace keys, TLS 1.3 in transit

AES-256 at rest with per-workspace keys, TLS 1.3 in transit

Access

Access

SAML SSO, SCIM provisioning, scoped API keys

SAML SSO, SCIM provisioning, scoped API keys

Audit

Audit

Every run, approval and config change, exportable to your SIEM

Every run, approval and config change, exportable to your SIEM

Self-hosting

Self-hosting

Run in your AWS, GCP or Azure account on Enterprise

Run in your AWS, GCP or Azure account on Enterprise

Payload redaction

Payload redaction

Mark fields as secret and they never leave your cloud

Mark fields as secret and they never leave your cloud

Uptime

Uptime

99.95% SLA on Business, 99.99% on Enterprise

99.95% SLA on Business, 99.99% on Enterprise

Get started

Ship the agent. Keep the receipts.

Free for 50,000 steps a month. No credit card, no separate workers, and your first durable workflow deployed before lunch.

orrindel

Durable runtime for AI agents and automations. Every run, in plain text.

Book a 20-minute demo →
All systems normal
99.99% uptime · last 90 days
90 days agotoday
regions us-east · eu-west · ap-south
soc 2 type ii · gdpr · hipaa (baa)
sdk v4.2.1 · node · python · go
© 2026 Orrindel Labs, Inc.PrivacyTermsSecurityMade in plain text.

Create a free website with Framer, the website builder loved by startups, designers and agencies.