Should AI investigate a production incident?
A worked architecture decision for read-only AI incident investigation, human command, evidence provenance, and safe failure during an outage.
What makes the obvious answer unsafe?
An incident assistant can correlate alerts, telemetry, and change records, but correlation is not a rollback instruction. A human incident commander still owns impact, mitigation, and communication.
Three credible approaches
Compare human-only investigation, deterministic evidence collection, and a read-only assistant that returns cited hypotheses and contradictions.
Where the authority changes hands
Use bounded retrieval and source-linked evidence. The assistant cannot access production write tools. A commander tests the lead and uses existing recovery controls.
What must remain true?
Preserve evidence provenance, uncertainty, least privilege, and a human-only fallback when retrieval or the model fails.
Test the decision, not just the diagram
Rehearse false leads, missing telemetry, malicious log content, denied writes, and incident-time latency.