SaaS incident response runbook: manage an outage and write a postmortem
Create a usable SaaS incident-response runbook with clear roles, customer updates, an event timeline, recovery steps and a blameless postmortem process.
In this guide
What is SaaS incident response?
Incident response is the coordinated work of limiting customer impact, restoring service, communicating clearly and learning from an unplanned event. Fixing a technical symptom and managing an incident are related but different tasks: responders need both a way to mitigate the fault and a way to organise people and information. Google SRE recommends declaring incidents early, keeping a working record, and assigning clear command, operations and communication responsibilities.
Write a trigger that makes it easy to declare an incident
Define practical signals such as a sustained service-level objective breach, suspected data exposure, widespread payment failure or a serious security alert. State who can declare an incident and where responders join. A clear trigger reduces debate while customers may be affected; severity can be adjusted as evidence improves.
Assign an incident commander, operations lead and communications lead
The incident commander coordinates priorities and decisions; the operations lead investigates and applies technical mitigations; the communications lead updates responders, stakeholders and customers. A small startup may have one person cover more than one role, but name the responsibility and hand it over explicitly when workload grows.
Keep one shared timeline and decision log
Record when the event began, alerts, affected services, customer impact, investigation results, actions, owners and update times. Distinguish observations from hypotheses and note when a hypothesis changes. This gives responders a reliable handoff and creates evidence for the later review.
| Time / observation | Customer or service impact | Action and owner | Decision / evidence | Next update |
|---|---|---|---|---|
What should an outage runbook tell responders to do?
Establish scope and stabilise the service
Check health dashboards, recent changes and affected workflows. Confirm whether impact is limited to one tenant, region, integration or customer segment. Prefer a reversible mitigation that reduces harm, such as pausing a risky job or rolling back a change, and record the expected effect before acting.
Protect customer data and evidence
For a suspected security or privacy event, limit further access, preserve relevant logs and involve the designated security and privacy contacts. Avoid running cleanup that destroys useful evidence. Determine which data, accounts and tenants may be involved; qualified advisers should guide contractual, legal and regulatory notification decisions.
Send useful, time-bound status updates
State what is affected, who may be impacted, known workarounds, what the team is doing and when the next update will appear. Use plain language and a stable public status page where available. If the cause is unknown, say so; do not speculate or promise a restoration deadline without support.
How do you write a useful postmortem after a SaaS outage?
Reconstruct the event from evidence
Build a timeline from alerts, logs, deploy records and the incident notes. Describe customer-visible impact, duration, detection, response and recovery. Separate confirmed facts from possible contributing factors, and explain where visibility was missing without turning the review into a search for an individual to blame.
Identify contributing conditions and missed safeguards
Ask why the change or failure reached customers, why it was not detected sooner and why existing controls did not limit the impact. More than one condition usually contributed. Review deploy safety, capacity, testing, alerts, access controls, documentation and handoffs where relevant.
Turn findings into a few owned, verifiable actions
Write corrective actions with a named owner, due date and evidence of completion. Prefer changes that prevent recurrence or detect it earlier, such as a tenant-isolation test, safer rollout, tested restore or clearer alert. Track actions to completion and share lessons with affected teams and customers at an appropriate level of detail.
SaaS incident-response questions
Should a small startup declare an incident before the cause is known?
Yes, when a defined customer-impact or safety trigger is met. The team can update severity and working hypotheses as it learns more. Declaring early gives people one place to coordinate and communicate.
What is the difference between an incident report and a postmortem?
An incident report is the live record used to coordinate response and handoffs. A postmortem is a later review of impact, timeline, contributing conditions and preventive actions. The report supplies evidence; it does not replace the analysis.
Should a postmortem name the person who made a mistake?
Focus on system conditions, decisions, information and safeguards rather than blame. A learning-oriented review makes it easier to report risks and improve controls. Handle individual conduct through the appropriate separate process if needed.
When should customers receive an incident update?
Use the communication cadence in the incident plan and update customers when verified impact or meaningful status changes. If an investigation takes time, publish the next update time so customers know when to check back. Follow any applicable contract or legal notification duties.
Related practical guides
Related issue guides
Sources and publication record
Draft prepared 27 September 2026; project-team editorial review pending · Sources checked .
- Google SRE Workbook: Incident responseGoogle Site Reliability Engineering
- Google Cloud: Architecting disaster recovery for cloud infrastructure outagesGoogle Cloud
- AWS Well-Architected: operational readiness reviewAmazon Web Services
- AWS Well-Architected: plan for unsuccessful changesAmazon Web Services