Autonomous incident investigation
It isn’t told what broke.
It finds out.
IncidentOS hands Claude a production incident — real telemetry, a real repository, a real defect — and no answer. It builds competing hypotheses, disproves all but one, patches the code, and verifies the fix against a real measurement.
Claude Opus 5 · 11 tools · deterministic incident data · no Kubernetes required
Investigation
waiting for the incident_
Hypotheses
none yet
A replay of an actual investigation. Every number in it is one the real run produces.
Ask an AI what might be wrong and it will tell you something plausible. Plausible is not the same as true.
So the investigation is structured so that a conclusion has to be earned. The agent records its findings through tools that refuse to accept a shortcut — and the measurements that verify the fix are taken by the server, not by the model that wrote it.
How it works
Three phases, one thread.
The agent keeps everything it has learned across all three, so the engineer writing the fix is the same one who proved the cause.
Investigate
It starts with nothing but the incident card. It pulls metrics, logs, Kubernetes events and deployment history, and builds a timeline of what actually happened — correlating across sources, because one signal is a coincidence.
get_metricssearch_logsget_kubernetes_eventsget_deployment_historyProve
It proposes competing explanations — including ones it expects to be wrong — and then tries to disprove each. Traffic rose only 8%. The connection pool never queued. Every finding cites the tool result it came from.
report_hypothesisreport_evidencesearch_coderead_fileFix and verify
Once the cause is proven it patches the source, type-checks it, runs the suite, and re-measures memory. If a test fails it reads the output and iterates. The numbers are measured, not asserted.
report_root_causeapply_patchrun_testsRigour, enforced
Rules the agent cannot talk its way around.
Prompts ask. These refuse. Each rule lives inside the tool that records the finding, so skipping a step returns an error the agent has to correct — not a timeline with holes in it.
- Competing hypotheses
- At least two recorded before any conclusion is accepted.
- Cited evidence
- Two supporting findings minimum, each naming the tool result behind it.
- Elimination
- Every hypothesis examined — including the ones being ruled out.
- Causality
- No patch can be applied before a root cause is proven.
Verification
Every number here was measured.
The demo repository contains a real memory leak and a real test that catches it. The before and after figures come from running that code — once on the defect, once on the patch the agent wrote.
Test suite
47 / 47
A regression test that fails while the defect is present, and passes after the patch.
Retained memory
3.55 MB
Measured over 25,000 real captures — heap plus external, which is what the cgroup counts.
Per transaction
149 B
The whole request was being retained where an 80-byte summary would do.
A note on that memory figure: the retained payloads are typed-array backing stores, which live outside the V8 heap. Measuring heapUsed alone reports a comfortable 14 MB for a process that is about to be OOM-killed — the same trap the incident itself is about.
Give it an incident.
One click, about ninety seconds, and you can watch every tool call it makes on the way to the answer.