The problem
Ingestion pipelines fail in the boring middle: a PDF that times out, an embedding batch that hits a rate limit, a job that dies at file 8,000 of 12,000 and starts again at file one.
How the workflow runs
Each uploaded file triggers its own run, with concurrency capped at 20 per workspace so the embeddings provider never sees a spike. Extraction, chunking and embedding are separate steps, so a rate-limited batch retries on its own while the extracted text stays saved.
A nightly cron run compares the index with the bucket and cleans up chunks for deleted files. Large files use step.sleep between batches instead of holding a worker open.
What changes
A failed file is a single red row in the dashboard with the exact page that broke, not a mystery gap in search results three weeks later.


