learn

Production Incident Checklist

A concise checklist for responding to cloud and application incidents without losing focus on service recovery.

Troubleshooting

Start with the symptom. End with a verified fix.

What’s wrong?

A production incident has customer or service impact.

Possible causes

  1. Recent change
  2. dependency failure
  3. infrastructure saturation
  4. configuration issue
  5. regional or workload-specific failure

Diagnosis

  1. Confirm impact
  2. check dashboards alerts logs and recent changes
  3. determine scope
  4. keep one hypothesis at a time

2 · Fix

Stop unsafe changes and use a reversible mitigation, rollback, or failover when appropriate.

Verify

Verify service recovery from the user's perspective and preserve evidence.

3 · Prevent

Identify root cause and contributing factors, then update the runbook and prevention actions.

Production Incident Checklist

First 10 minutes

  • Confirm the incident and affected service.
  • Identify customer impact.
  • Assign an incident owner.
  • Check dashboards, alerts, logs, and recent changes.
  • Determine whether the issue is isolated or widespread.
  • Stop unsafe changes and consider rollback or failover.

During mitigation

  • Keep one clear hypothesis at a time.
  • Record important observations and actions.
  • Prefer reversible changes.
  • Communicate impact and status clearly.
  • Do not confuse a workaround with root cause.

After recovery

  • Verify the service from the user's perspective.
  • Preserve relevant logs and evidence.
  • Identify root cause and contributing factors.
  • Create prevention actions with owners.
  • Update the runbook so the next incident is faster to resolve.

Learning checkpoint

Mark this guide complete to include it in your local Engineering Journey.

Knowledge path

Connected concepts

Explore the knowledge graph

WATCH WITH THIS TOPIC