5 Copy Ready Runbook Examples and Reusable Template for Ops Teams

A runbook is a compact, action-oriented procedure that specifies a trigger condition, an ordered sequence of diagnostic and remediation steps, verification commands, rollback instructions, and an escalation path. This article supplies five copy-adaptable runbook examples, a reusable structural template, and maintenance recommendations grounded in AWS Well-Architected guidance, NIST incident handling practice, and conventions drawn from site reliability engineering documentation, published here.
TL;DR:
Runbooks should include explicit rollback steps, especially for critical procedures like database failovers or system patches, to prevent unresolved issues.
Storing runbooks in the same code repository as infrastructure ensures synchronization between documentation and system state, while redundant offline copies protect against outages.
Regular testing, including quarterly tabletop exercises and biannual full run-throughs, is essential to confirm procedures remain accurate and effective under pressure.
A clear ownership, last-updated date, and verification commands are vital to keep runbooks reliable and prevent drift or outdated procedures during incidents.
Using a structured, executable format enables both automation and human understanding, reducing errors and improving response consistency.
PROJECT-JTH
Bring More Clarity Under Pressure
PROJECT-JTH provides technical advisory services that help teams address operational challenges and maintain effectiveness in mission-critical environments.
Table of Contents
1. Five runbook examples you can adapt today
Each example below follows the same skeleton: purpose, trigger, ordered steps, verification, rollback, and escalation contact format. The steps are deliberately generic so they translate across most infrastructure stacks.
Monitor folder or log directory near capacity. Trigger: disk utilization on the monitored volume exceeds a defined threshold for 15 minutes. Steps: confirm alert with
df -h, identify the largest consumers withdu -sh /var/log/* | sort -rh, archive or compress logs older than 30 days, rotate active logs, confirm space recovered withdf -h, and notify the service owner if usage remains above threshold. Rollback: restore archived files from backup if a required log was deleted in error. Escalation: on-call storage engineer, then infrastructure lead if unresolved in 30 minutes.Database failover. Trigger: primary database health check fails three consecutive times. Steps: verify replica lag is below an acceptable window, promote the standby with your orchestration tool, update the connection string or DNS record, confirm application writes succeed against the new primary, and monitor error rates for 10 minutes. Verification: run a test transaction and confirm it commits. Rollback: demote the promoted node and restore the original primary if data integrity checks fail. Escalation: database on-call, then engineering manager.
TLS certificate renewal. Trigger: certificate expiration within 14 days, detected by a monitoring check. Steps: generate a certificate signing request with
openssl req -new -key domain.key -out domain.csr, submit to the certificate authority, retrieve the signed certificate, deploy it to the load balancer or web server, reload the service withsystemctl reload nginx, and verify the new expiration date withopenssl x509 -enddate -noout -in domain.crt. Rollback: restore the prior certificate file and reload. Escalation: security engineer on call.Server patching. Trigger: scheduled maintenance window or a critical security advisory. Steps: drain traffic from the node, snapshot the current state, apply patches, reboot if required, run smoke tests against core services, and return the node to the load balancer pool. Verification: confirm service health checks pass for five consecutive minutes. Rollback: restore from the pre-patch snapshot. Escalation: platform on-call, then change advisory board contact.
New-employee provisioning. Trigger: HR system generates a new-hire ticket. Steps: create identity provider account, assign role-based access groups, provision hardware or virtual desktop, issue credentials through the secure channel, confirm access to required systems, and document completion in the ticketing system. Verification: new hire successfully logs into at least one core system. Rollback: revoke access immediately if provisioned incorrectly. Escalation: IT service desk lead.
2. Build a reusable runbook template and structure
A single enforced skeleton keeps every runbook usable under pressure, regardless of who wrote it or when.
Metadata: owner, last-updated date, and system name, so staleness is detectable at a glance.
Purpose and desired outcome: the specific result the runbook restores or achieves.
Trigger or alert signature: the exact condition, ideally a numeric threshold, that starts the procedure.
Prechecks and diagnostics: commands that confirm the problem before action is taken.
Step-by-step actions: the ordered remediation sequence, numbered for execution under stress.
Validation and verification: commands or checks proving the fix worked.
Rollback and abort conditions: how to reverse the action if it fails or makes things worse.
Escalation contacts: named roles, not individuals, with a clear time boundary for handoff.
Related documentation and links: pointers to architecture diagrams or dependent runbooks.
Post-incident actions: follow-up tasks such as ticket closure or retrospective notes.
These fields align closely with the components AWS Well-Architected identifies as core to a repeatable runbook: outcome, procedure, validation, and escalation are treated as non-negotiable rather than optional.
Pro tip: Keep each runbook focused on a single failure mode, and link out to a higher-level playbook when an incident spans multiple systems.

3. Choose an executable format and store runbooks safely
Executable formats such as runbook.md or the .runbook file convention embed typed blocks (check, step, rollback, wait) alongside plain-language instructions, letting a script and a human read the same document. The runbook file format supports frontmatter metadata and structured blocks that keep documentation and commands from drifting apart over time.
Store runbooks in the same Git repository as the service code they describe, so changes to infrastructure and changes to procedure ship together.
Cross-link runbooks from alerting rules and dashboards so the on-call engineer reaches the correct document within seconds.
Mirror critical runbooks to redundant, cloud-independent storage and keep an offline exported copy for on-call staff, since the primary hosting platform may itself be part of the outage.
4. Test and maintain runbooks on a fixed schedule
A runbook that has not been exercised recently should be treated as unverified, not trusted. NIST’s incident handling guidance recommends periodic exercise of response procedures so they perform correctly during genuine incidents rather than only in theory.
Run a smoke test immediately after any change to the underlying system or the runbook itself.
Schedule a quarterly tabletop exercise where the team walks through the steps without executing them against production.
Perform a full run-through against a staging environment at least twice a year, timing mean time to restore and confirming every verification check passes.
Require sign-off from an observer during drills, and route runbook edits through pull request review with linting on executable sections.
Pro tip: Tag each runbook with its last successful drill date so stale procedures surface automatically instead of being discovered mid-incident.
5. Avoid these common runbook anti-patterns
Storing runbooks on infrastructure that might be down during the outage they are meant to resolve, which AWS identifies as the most dangerous anti-pattern in operational documentation.
Missing rollback steps, leaving an operator unable to reverse a failed remediation.
Oversized, multi-system documents that force a reader to skim past irrelevant sections during a crisis.
No named owner or last-updated field, which allows procedures to drift silently out of accuracy.
Runbooks that are never exercised, so the first real test happens during an actual outage.
A short audit checklist covers ownership, storage redundancy, rollback presence, verification commands, escalation contacts, and drill recency.
6. Why disciplined runbooks matter under pressure
Teams that enforce a single runbook structure and drill it regularly tend to restore service faster and hand off incidents with less confusion, because everyone is reading from the same skeleton rather than improvising. Technical advisory work that pairs infrastructure review with communication training for high-pressure environments treats runbook discipline as a leading indicator of operational maturity. A focused runbook audit is often the fastest way to surface where documentation and reality have drifted apart.
— Jesse Hart
7. How PROJECT-JTH can help you build and audit runbooks
Teams that recognize gaps in their own runbooks, missing rollbacks, undocumented owners, procedures nobody has tested since deployment, don’t have to close them alone. Certain technical systems and risk reviews and infrastructure automation reviews can examine existing procedures and automation for weaknesses, and dedicated software gives teams a central place to store production records and runbook history instead of scattering them across chat threads and personal drives.

A runbook without a tested rollback path is a liability disguised as documentation. That single gap accounts for a large share of failed recovery attempts across the anti-patterns AWS flags in its operational guidance.
Starting point | What it addresses |
|---|---|
Surfaces missing owners, stale procedures, and untested rollback steps | |
Infrastructure Automation & Ansible Review | Converts manual runbook steps into automated, repeatable actions |
Centralizes production records and runbook artifacts in one accessible system |
Teams ready to close these gaps can start with a technical consulting engagement to audit existing runbooks or automate the ones that matter most.
Sources
For deeper reference, AWS Well-Architected’s runbook guidance covers core structure and anti-patterns, NIST SP 800-61r3 addresses testing and maintenance, the AIOps SRE runbook template offers a downloadable production-ready skeleton, and the runbook file format documentation details executable block syntax.
FAQ
What is an example of a runbook?
A database failover runbook is a common example: it defines a health check trigger, steps to promote a standby database, a command to verify the new primary accepts writes, and a rollback path if the promotion fails. Monitoring, patching, and certificate renewal procedures follow the same pattern with different steps.
How do you write a runbook?
Start from a fixed template that includes metadata, a precise trigger, ordered steps, verification commands, rollback instructions, and escalation contacts, then fill in the specifics for one failure mode at a time. Test the draft against a staging environment before trusting it in production, since an untested runbook is not a reliable one.
What exactly is a runbook?
A runbook is a documented, step-by-step procedure built to produce a consistent, repeatable outcome for a specific operational event, as described in AWS Well-Architected guidance. It differs from a playbook, which typically coordinates multiple runbooks across a larger incident.
What are the different types of runbooks?
Common types include monitoring and alert response, failover and disaster recovery, patching and maintenance, certificate and credential renewal, and provisioning or deprovisioning procedures. Each type shares the same underlying structure of trigger, steps, verification, rollback, and escalation, even though the specific actions differ.
