The problem
The first ten minutes of an incident are spent finding the same five pieces of context: what changed, what’s erroring, which customers are affected and who owns it.
How the workflow runs
An alert triggers the run. Tool steps pull the last three deploys, the error rate by service and a sample of matching log lines in parallel. An LLM step writes a five-line summary with the most likely cause and links to the evidence.
If the alert resolves on its own within five minutes, waitForEvent catches the recovery and the run posts a quiet note instead of paging anyone.
What changes
The on-call engineer opens the page to a hypothesis and the evidence behind it, not a raw graph.



