← Back to shipped cases

01 / SHIPPED CASE

AI-SRE: an incident partner that has to show its work

The LLM helps form and rank hypotheses. Read-only evidence establishes system state. Deterministic policy and a human gate decide whether any bounded remediation may run.

PUBLIC CASE STUDY / Names, environments, endpoints, and operational details are generalised. The 27–30 second range describes three isolated complex rollback drills—not general MTTR or a production SLA.

≈25SCENARIOS / 2 REPOS / 2 CLUSTERS
8REAL ENGINEERING GAPS FOUND + FIXED
0DANGEROUS ACTIONS IN CONTROLLED SAFETY TESTS
27–30sCOMPLEX ROLLBACK DRILL / N=3
ROLESystem design · implementation · validation
TIME2026.06 — 2026.07
COLLABORATIONPlatform / SRE · on-call approval path
SCOPENonprod shadow + isolated chaos validation

THE PROBLEM

Useful autonomy begins with accountable decisions—not fewer people.

Incident work is often a context problem: what changed, where the failure began, what is affected, and what can be checked safely next.

AI can accelerate that reasoning, but it is not production truth and does not own authority. That separation is the core of this system.

SANITISED WALKTHROUGH

A bad image rollout, traced from signal to verified recovery.

This reconstruction follows project test and audit records. Service, namespace, repository, and internal communication details are removed.

INCIDENT / REDACTEDRECOVERED
01 / INPUTA new revision never becomes ready

isolated nonprod chaos namespace

02 / REASONRank a bad-image regression as the leading hypothesis

reasoning does not confer authority

03 / EVIDENCENew revision unhealthy; previous revision was healthy

status + rollout history

04 / GATESTarget, scope, and allowlist checks pass

the sandbox drill may auto-approve; production still requires a person

05 / RESULTRollback completes and the post-check returns Ready

27–30s across three complex drills

SANITISED AUDIT TRAIL

The retained record covers evidence, policy, approval, execution, and the post-check—not just model prose.

candidate_action: rollback_deployment
evidence: [new_revision_unhealthy, previous_revision_healthy]
policy: allowed_nonprod_sandbox
approval: sandbox_drill_only
post_check: recovered
audit_status: complete

CONTROL MODEL

The LLM explains risk. Deterministic controls own permission.

LLM

Hypotheses and prioritisation

Proposes what may be happening and which read-only check should come next.

EVIDENCE

System state, not confidence theatre

Important claims point back to observed signals; missing evidence lowers confidence.

DETERMINISTIC

Policy, scope, denylist, signatures

Validation failures close the path instead of asking the model to be more careful.

HUMAN

Explicit production approval

The operator sees evidence, blast radius, and rollback plan before approving or rejecting.

WHAT THE EVIDENCE SUPPORTS

The engineering loop is validated. User outcomes still need a publishable baseline.

≈25

Standard and complex scenarios across two repositories and two clusters.

8

Real engineering gaps found and fixed during validation.

0

Dangerous actions in controlled safety tests.

OPEN

Time saved, steps removed, production adoption, and user feedback are not publicly claimed yet.

REQUEST A WALKTHROUGH

Want the architecture, a sanitised demo, or a conversation about a similar workflow?

Email Ian ↗

NEXT SHIPPED CASE

How the same evidence boundary became an engineering knowledge agent.

Read RelayOps ↗