← Back to projects
SREActiveOpen source

SRE Incident-Response Runbooks

The operational knowledge that keeps services reliable: runbooks (crashloop, high error rate, CPU, DB exhaustion), an SLI/SLO/error-budget framework with burn-rate alerting, and postmortem + incident-comms templates.

Highlights

  • Actionable 3am-ready runbooks
  • SLO + error-budget framework
  • Blameless postmortem templates

Tech Stack

SRESLOIncident ResponsePostmortemsReliability

Reactions & comments