Alert Runbooks
Alert Runbooks
What to do when monitoring fires — separate the real condition from the noise and act on the right one.
What to do when monitoring fires — separate the real condition from the noise and act on the right one.
AV/EDR Agent Offline Alert
Triage an AV/EDR agent-offline alert: decide if the device is off or up with a dead agent, quantify unprotected time, and route on the protection gap.
Backup Missed vs Failed Alert
Distinguish a backup that never ran (missed) from one that ran and errored (failed) (two different routes), and always state exposure via last-known-good.
Certificate Expiry Alert
Triage a certificate expiry alert: tier urgency by days remaining, identify what the cert secures and who owns renewal, and route into renewal work.
Disk Space Alert
Triage a low-disk-space alert from any monitor: separate threshold noise from real pressure, read growth rate from history, rank consumer hypotheses.
High CPU/Memory Alert
Triage a CPU or memory threshold alert: separate a transient spike from sustained pressure via history, and route servers versus workstations differently.
Patch Failure Alert
Triage a patch-failure alert: separate a one-off from a repeat offender, detect reboot-pending as the usual culprit, correlate against the patch window.
RAID Degradation Alert
Triage a RAID degraded or failed-member alert with zero-margin urgency (one failure from data loss), and enforce the verify-backups-BEFORE-rebuild rule.
Was this page helpful?