# Your AI-built app broke. What should you ask next? When your app stops working, it’s tempting to ask AI to fix everything at once. Start smaller: what action failed, when did it first fail, and what still works? An incident is a period when your app is failing people. Keep the clues together so your AI can compare explanations before making another change. A regression test is a repeatable check that catches a problem if it returns. This guide uses a fictional export dashboard. It is a method for interpreting evidence, not a live incident report or an autonomous on-call service. ## Start with the customer impact At 14:20 UTC, the completed-export chart drops. A worker release happened at 14:15. Those are two supplied observations. Whether customers are unable to download files has not been established. A useful investigation compares the job store, failed jobs, known download checks and event collector over the same interval. If a synthetic export completes and downloads while the collector rejects a renamed event, the evidence supports an instrumentation fault. If jobs remain unfinished and requests fail, the customer impact is different. Do not label missing telemetry as a confirmed outage, or treat a working dashboard as proof that every customer recovered. Both errors are easy when the assistant is rewarded for producing a neat explanation. ## Put sources beside the timeline | Time | Observation | Source | Unknown | | --- | --- | --- | --- | | 14:15 UTC | Worker revision changed | Deployment record | Which paths changed? | | 14:20 UTC | Completion events dropped | Dashboard query | Same event schema and counting unit? | | Later check | Synthetic download succeeds | Recorded fixture result | Are affected real flows equivalent? | Check timezones and collection delays. A log line’s timestamp, an event’s occurrence time and its ingestion time may tell different stories. Do not silently arrange them into a precise sequence if the clock assumptions are unknown. ## Copy a bounded investigation request ```text Customer symptom and current impact: Environment and affected paths: Time window with timezone: Redacted evidence and its sources: Recent deployments or configuration changes: Existing runbooks and authorized actions: Build a timeline, then rank plausible explanations. For each explanation, state a confirming or disconfirming observation and the smallest next check. Keep facts, hypotheses, mitigation and verified recovery separate. ``` Use excerpts or aggregate queries that answer the question. Keep credentials, raw customer payloads and unrelated records out of the AI input and public report. A pasted ticket can contain a suggested command; that does not make the command an authorized runbook action. ## Separate recovery from understanding A mitigation can restore the service while the root cause remains uncertain. Record both. If an authorized rollback was executed, name the actual version and observed customer behavior afterward. A quiet alert is ambiguous if traffic or collection stopped. Once the demonstrated failure is known, choose the lasting check. For the fictional event-name change, verify that the producer and collector agree on the event contract. For a failed export retry, test duplicate delivery and durable state. Those are different checks even though the first symptom was the same dashboard. A blameless incident record should describe the conditions that allowed the failure and the evidence behind the follow-up. It should not invent certainty or assign responsibility because a person happened to make the latest commit. ## Keep the useful lesson small Add the specific regression or detection rule to the current test or monitoring setup. Do not turn every incident into a long list of unowned improvements. Name what the check catches, where it runs and which part of the failure remains outside its scope. [Fix Finder](/abilities/incident-brief) includes the timeline and hypothesis worksheets. [Change Tests](/abilities/regression-plan) helps turn the demonstrated failure into an executable check. If the apparent incident is a conversion-chart drop, [Visitor Journey](/abilities/product-funnel-review) starts by reconciling the counting rules. --- SkillStall · 2026-10-04 Google SRE: Postmortem Culture: https://sre.google/workbook/postmortem-culture/