By·

How to Write Incident Postmortems with Mermaid Diagrams (RCA That Engineers Actually Read)

Use Mermaid sequence diagrams, timelines, and flowcharts to document incident root cause analysis, blast radius, and remediation — with copy-paste templates for SRE and DevOps teams.

Rendered How to Write Incident Postmortems with Mermaid Diagrams (RCA That Engineers Actually Read)
Rendered example. Copy the first code block below to edit it.

# How to Write Incident Postmortems with Mermaid Diagrams (RCA That Engineers Actually Read)

Postmortems have a reputation problem. Most are sprawling Google Docs that nobody reads after the post-incident review meeting. The root cause gets buried in paragraph six of a twelve-paragraph narrative, and the timeline — the single most important artifact — is a bulleted list that requires mental gymnastics to reconstruct.

Mermaid fixes this. A well-placed sequence diagram, timeline, or flowchart shows the entire incident in one glance. Engineers scan it in seconds. Stakeholders understand it without asking "so what actually happened?"

This guide gives you copy-paste templates for the four diagram types every solid postmortem needs.

---

Why Diagrams in Postmortems?

A postmortem answers three questions. Each maps to a diagram type:

QuestionBest DiagramWhat It Shows
What happened and when?TimelineChronological sequence of events
Which systems talked to each other?Sequence DiagramCross-service interactions
What was the blast radius?FlowchartAffected components and cascading failures
What fixed it (and what prevents it)?FlowchartRemediation path + guardrails

Text answers these questions too. Diagrams make them *immediate*.

---

1. Incident Timeline

The timeline is the backbone. Every postmortem needs one. Here's a real-looking incident — a bad config deploy that took down payments:

timeline
    title Payment Service Outage — Oct 1 2026
    section Pre-Incident
      14:22 UTC : Config PR #1427 merged
      14:30 UTC : CI/CD pipeline deploys to staging
      14:35 UTC : Staging smoke tests PASS (no payment checkout test)
    section Incident
      15:00 UTC : Deploy to production begins
      15:02 UTC : Payment API 500 errors begin (Sentry spike)
      15:04 UTC : PagerDuty alert fires — on-call engineer ack
      15:08 UTC : Engineer identifies config change as cause
      15:11 UTC : Rollback initiated
      15:14 UTC : Rollback complete — errors drop to 0
    section Post-Incident
      15:20 UTC : Postmortem doc created
      15:45 UTC : Hotfix PR opened (add env var validation)
      16:30 UTC : Payment checkout added to staging smoke suite
Try in Editor →

The timeline tells the story in 15 seconds. Everyone — from the CTO to the junior SRE — can follow it.

---

2. System Interaction Diagram

What actually broke? A sequence diagram shows the exact request path that failed:

sequenceDiagram
    participant User
    participant Gateway as API Gateway
    participant Payment as Payment Service
    participant Config as Config Service
    participant Stripe

    User->>Gateway: POST /checkout { cart }
    Gateway->>Payment: Forward checkout request
    Payment->>Config: GET payment.stripe_api_key
    Config-->>Payment: null ❌ (env var unset after bad config deploy)
    Payment--xGateway: 500 Internal Server Error
    Gateway--xUser: 500 — Checkout failed
Try in Editor →

One diagram, root cause visible immediately: the Config Service returned null for a required key, and the Payment Service had no fallback or validation.

---

3. Blast Radius / Cascading Failure Map

What else got affected? A flowchart maps the dependency graph of the failure:

flowchart TD
    A["Config Deploy<br/>(missing STRIPE_API_KEY)"] --> B["Payment Service<br/>❌ 500 errors"]
    B --> C["Checkout Page<br/>❌ Unavailable"]
    B --> D["Subscription Renewals<br/>❌ Payment failures"]
    D --> E["Billing Alerts<br/>⚠️ Delayed invoices"]
    C --> F["Revenue Impact<br/>💰 ~$12K lost bookings"]
    
    A --> G["Config Service<br/>⚠️ No validation on deploy"]
    
    style A fill:#f96,stroke:#333
    style B fill:#f96,stroke:#333
    style C fill:#f96,stroke:#333
    style D fill:#f9c,stroke:#333
    style E fill:#ff9,stroke:#333
    style F fill:#f66,stroke:#333
    style G fill:#fc9,stroke:#333
Try in Editor →

The red/pink nodes show direct impact. Yellow shows indirect. This diagram replaces three paragraphs of prose about "cascading effects."

---

4. Remediation & Prevention Flow

What's the fix, and what stops it next time?

flowchart LR
    subgraph Immediate["🚨 Immediate Fix"]
        A["Rollback config deploy"] --> B["Payment service recovers"]
    end
    
    subgraph Short["🔧 Short-Term (Same Day)"]
        C["Add env var validation<br/>on Payment Service startup"] --> D["Service refuses to start<br/>if required keys missing"]
    end
    
    subgraph Long["🛡️ Long-Term Prevention"]
        E["Add payment checkout<br/>to staging smoke suite"] --> F["Catch missing keys<br/>before production deploy"]
        G["Config Service: pre-deploy<br/>schema validation"] --> H["Never deploy<br/>incomplete configs"]
    end
    
    Immediate --> Short --> Long
Try in Editor →

This shows the graduated response: stop the bleeding → harden the service → add systemic guardrails. Action items map directly to boxes on this chart.

---

5. Full Postmortem Template

Combine everything into a single markdown doc:

# Incident Postmortem: Payment Service Outage
**Date:** 2026-10-01 | **Duration:** 12 minutes | **Severity:** SEV2

## Timeline
> [Insert Mermaid timeline from Section 1]

## Root Cause
> [Insert sequence diagram from Section 2]

> **Description:** Config deploy removed STRIPE_API_KEY environment variable. 
> Payment Service had no startup validation for required config keys.

## Blast Radius
> [Insert flowchart from Section 3]

> **Impact:** Checkout unavailable for 12 min. ~180 affected users. ~$12K estimated revenue impact.

## Remediation
> [Insert flowchart from Section 4]

## Action Items
| # | Action | Owner | Due |
|---|--------|-------|-----|
| 1 | Add env var validation on Payment Service startup | @sre-team | Oct 2 |
| 2 | Add checkout flow to staging smoke suite | @qa | Oct 3 |
| 3 | Config Service: pre-deploy schema validation | @platform | Oct 10 |

## Lessons Learned
- Required config keys must be validated at service startup, not assumed present.
- Staging smoke tests must cover the critical revenue path.
- Rollback was fast (2 min) — CI/CD pipeline performed well under pressure.

---

Why Mermaid Wins for Postmortems

AlternativeProblem
Google Slides / VisioBinary files, version-control nightmares, can't copy-paste into markdown docs
Screenshots of drawn diagramsUnsearchable, uneditable, blurry in dark mode
Long-form text onlyNobody reads them. Seriously.
MermaidVersion-controlled in git right next to docs. Searchable. Editable. Renders in GitHub, GitLab, Notion, Obsidian, and Confluence.

---

Pro Tips

  1. Always include a timeline. Even a simple one. It anchors the entire narrative.
  2. Color-code severity in flowcharts: red for direct impact, yellow for indirect, green for unaffected. Stakeholders scan colors, not labels.
  3. Keep sequence diagrams to 5–6 participants max. More than that and you lose the "glanceable" quality.
  4. Version your postmortems alongside your code. When the same incident happens again (and it will), you can diff what changed.
  5. Use Note over annotations in sequence diagrams to mark the exact moment of failure, not just the arrows.

The best postmortem is the one your team actually reads six months later. Mermaid makes that happen.