How to Write Incident Postmortems with Mermaid Diagrams (RCA That Engineers Actually Read)
Use Mermaid sequence diagrams, timelines, and flowcharts to document incident root cause analysis, blast radius, and remediation — with copy-paste templates for SRE and DevOps teams.
# How to Write Incident Postmortems with Mermaid Diagrams (RCA That Engineers Actually Read)
Postmortems have a reputation problem. Most are sprawling Google Docs that nobody reads after the post-incident review meeting. The root cause gets buried in paragraph six of a twelve-paragraph narrative, and the timeline — the single most important artifact — is a bulleted list that requires mental gymnastics to reconstruct.
Mermaid fixes this. A well-placed sequence diagram, timeline, or flowchart shows the entire incident in one glance. Engineers scan it in seconds. Stakeholders understand it without asking "so what actually happened?"
This guide gives you copy-paste templates for the four diagram types every solid postmortem needs.
---
Why Diagrams in Postmortems?
A postmortem answers three questions. Each maps to a diagram type:
| Question | Best Diagram | What It Shows |
|---|---|---|
| What happened and when? | Timeline | Chronological sequence of events |
| Which systems talked to each other? | Sequence Diagram | Cross-service interactions |
| What was the blast radius? | Flowchart | Affected components and cascading failures |
| What fixed it (and what prevents it)? | Flowchart | Remediation path + guardrails |
Text answers these questions too. Diagrams make them *immediate*.
---
1. Incident Timeline
The timeline is the backbone. Every postmortem needs one. Here's a real-looking incident — a bad config deploy that took down payments:
timeline
title Payment Service Outage — Oct 1 2026
section Pre-Incident
14:22 UTC : Config PR #1427 merged
14:30 UTC : CI/CD pipeline deploys to staging
14:35 UTC : Staging smoke tests PASS (no payment checkout test)
section Incident
15:00 UTC : Deploy to production begins
15:02 UTC : Payment API 500 errors begin (Sentry spike)
15:04 UTC : PagerDuty alert fires — on-call engineer ack
15:08 UTC : Engineer identifies config change as cause
15:11 UTC : Rollback initiated
15:14 UTC : Rollback complete — errors drop to 0
section Post-Incident
15:20 UTC : Postmortem doc created
15:45 UTC : Hotfix PR opened (add env var validation)
16:30 UTC : Payment checkout added to staging smoke suiteTry in Editor →The timeline tells the story in 15 seconds. Everyone — from the CTO to the junior SRE — can follow it.
---
2. System Interaction Diagram
What actually broke? A sequence diagram shows the exact request path that failed:
sequenceDiagram
participant User
participant Gateway as API Gateway
participant Payment as Payment Service
participant Config as Config Service
participant Stripe
User->>Gateway: POST /checkout { cart }
Gateway->>Payment: Forward checkout request
Payment->>Config: GET payment.stripe_api_key
Config-->>Payment: null ❌ (env var unset after bad config deploy)
Payment--xGateway: 500 Internal Server Error
Gateway--xUser: 500 — Checkout failedTry in Editor →One diagram, root cause visible immediately: the Config Service returned null for a required key, and the Payment Service had no fallback or validation.
---
3. Blast Radius / Cascading Failure Map
What else got affected? A flowchart maps the dependency graph of the failure:
flowchart TD
A["Config Deploy<br/>(missing STRIPE_API_KEY)"] --> B["Payment Service<br/>❌ 500 errors"]
B --> C["Checkout Page<br/>❌ Unavailable"]
B --> D["Subscription Renewals<br/>❌ Payment failures"]
D --> E["Billing Alerts<br/>⚠️ Delayed invoices"]
C --> F["Revenue Impact<br/>💰 ~$12K lost bookings"]
A --> G["Config Service<br/>⚠️ No validation on deploy"]
style A fill:#f96,stroke:#333
style B fill:#f96,stroke:#333
style C fill:#f96,stroke:#333
style D fill:#f9c,stroke:#333
style E fill:#ff9,stroke:#333
style F fill:#f66,stroke:#333
style G fill:#fc9,stroke:#333Try in Editor →The red/pink nodes show direct impact. Yellow shows indirect. This diagram replaces three paragraphs of prose about "cascading effects."
---
4. Remediation & Prevention Flow
What's the fix, and what stops it next time?
flowchart LR
subgraph Immediate["🚨 Immediate Fix"]
A["Rollback config deploy"] --> B["Payment service recovers"]
end
subgraph Short["🔧 Short-Term (Same Day)"]
C["Add env var validation<br/>on Payment Service startup"] --> D["Service refuses to start<br/>if required keys missing"]
end
subgraph Long["🛡️ Long-Term Prevention"]
E["Add payment checkout<br/>to staging smoke suite"] --> F["Catch missing keys<br/>before production deploy"]
G["Config Service: pre-deploy<br/>schema validation"] --> H["Never deploy<br/>incomplete configs"]
end
Immediate --> Short --> LongTry in Editor →This shows the graduated response: stop the bleeding → harden the service → add systemic guardrails. Action items map directly to boxes on this chart.
---
5. Full Postmortem Template
Combine everything into a single markdown doc:
# Incident Postmortem: Payment Service Outage
**Date:** 2026-10-01 | **Duration:** 12 minutes | **Severity:** SEV2
## Timeline
> [Insert Mermaid timeline from Section 1]
## Root Cause
> [Insert sequence diagram from Section 2]
> **Description:** Config deploy removed STRIPE_API_KEY environment variable.
> Payment Service had no startup validation for required config keys.
## Blast Radius
> [Insert flowchart from Section 3]
> **Impact:** Checkout unavailable for 12 min. ~180 affected users. ~$12K estimated revenue impact.
## Remediation
> [Insert flowchart from Section 4]
## Action Items
| # | Action | Owner | Due |
|---|--------|-------|-----|
| 1 | Add env var validation on Payment Service startup | @sre-team | Oct 2 |
| 2 | Add checkout flow to staging smoke suite | @qa | Oct 3 |
| 3 | Config Service: pre-deploy schema validation | @platform | Oct 10 |
## Lessons Learned
- Required config keys must be validated at service startup, not assumed present.
- Staging smoke tests must cover the critical revenue path.
- Rollback was fast (2 min) — CI/CD pipeline performed well under pressure.---
Why Mermaid Wins for Postmortems
| Alternative | Problem |
|---|---|
| Google Slides / Visio | Binary files, version-control nightmares, can't copy-paste into markdown docs |
| Screenshots of drawn diagrams | Unsearchable, uneditable, blurry in dark mode |
| Long-form text only | Nobody reads them. Seriously. |
| Mermaid | Version-controlled in git right next to docs. Searchable. Editable. Renders in GitHub, GitLab, Notion, Obsidian, and Confluence. |
---
Pro Tips
- Always include a timeline. Even a simple one. It anchors the entire narrative.
- Color-code severity in flowcharts: red for direct impact, yellow for indirect, green for unaffected. Stakeholders scan colors, not labels.
- Keep sequence diagrams to 5–6 participants max. More than that and you lose the "glanceable" quality.
- Version your postmortems alongside your code. When the same incident happens again (and it will), you can diff what changed.
- Use
Note overannotations in sequence diagrams to mark the exact moment of failure, not just the arrows.
The best postmortem is the one your team actually reads six months later. Mermaid makes that happen.