Autonomous incident investigation

It isn’t told what broke.
It finds out.

IncidentOS hands Claude a production incident — real telemetry, a real repository, a real defect — and no answer. It builds competing hypotheses, disproves all but one, patches the code, and verifies the fix against a real measurement.

Claude Opus 5 · 11 tools · deterministic incident data · no Kubernetes required

INC-4821 · payments-api · SEV-1

Investigation

waiting for the incident_

Hypotheses

none yet

    INVESTIGATING…replay of a real run

    A replay of an actual investigation. Every number in it is one the real run produces.

    Ask an AI what might be wrong and it will tell you something plausible. Plausible is not the same as true.

    So the investigation is structured so that a conclusion has to be earned. The agent records its findings through tools that refuse to accept a shortcut — and the measurements that verify the fix are taken by the server, not by the model that wrote it.

    How it works

    Three phases, one thread.

    The agent keeps everything it has learned across all three, so the engineer writing the fix is the same one who proved the cause.

    01

    Investigate

    It starts with nothing but the incident card. It pulls metrics, logs, Kubernetes events and deployment history, and builds a timeline of what actually happened — correlating across sources, because one signal is a coincidence.

    get_metricssearch_logsget_kubernetes_eventsget_deployment_history
    02

    Prove

    It proposes competing explanations — including ones it expects to be wrong — and then tries to disprove each. Traffic rose only 8%. The connection pool never queued. Every finding cites the tool result it came from.

    report_hypothesisreport_evidencesearch_coderead_file
    03

    Fix and verify

    Once the cause is proven it patches the source, type-checks it, runs the suite, and re-measures memory. If a test fails it reads the output and iterates. The numbers are measured, not asserted.

    report_root_causeapply_patchrun_tests

    Rigour, enforced

    Rules the agent cannot talk its way around.

    Prompts ask. These refuse. Each rule lives inside the tool that records the finding, so skipping a step returns an error the agent has to correct — not a timeline with holes in it.

    Competing hypotheses
    At least two recorded before any conclusion is accepted.
    Cited evidence
    Two supporting findings minimum, each naming the tool result behind it.
    Elimination
    Every hypothesis examined — including the ones being ruled out.
    Causality
    No patch can be applied before a root cause is proven.

    Verification

    Every number here was measured.

    The demo repository contains a real memory leak and a real test that catches it. The before and after figures come from running that code — once on the defect, once on the patch the agent wrote.

    Test suite

    46 / 47→

    47 / 47

    A regression test that fails while the defect is present, and passes after the patch.

    Retained memory

    112.46 MB→

    3.55 MB

    Measured over 25,000 real captures — heap plus external, which is what the cgroup counts.

    Per transaction

    4,717 B→

    149 B

    The whole request was being retained where an 80-byte summary would do.

    A note on that memory figure: the retained payloads are typed-array backing stores, which live outside the V8 heap. Measuring heapUsed alone reports a comfortable 14 MB for a process that is about to be OOM-killed — the same trap the incident itself is about.

    Give it an incident.

    One click, about ninety seconds, and you can watch every tool call it makes on the way to the answer.