Alert Runbooks
Alert Runbooks
What to do when monitoring fires — separate the real condition from the noise and act on the right one.
What to do when monitoring fires — separate the real condition from the noise and act on the right one.
AV/EDR Agent Offline Alert
Triage an AV/EDR agent-offline alert — decide if the device is off or up with a dead agent, quantify unprotected time, and route on the protection gap.
Backup Missed vs Failed Alert
Distinguish a backup that never ran (missed) from one that ran and errored (failed) — two different routes — and always state exposure via last-known-good.
Certificate Expiry Alert
Triage a certificate expiry alert — tier urgency by days remaining, identify what the cert secures and who owns renewal, and route into renewal work.
Disk Space Alert
Triage a low-disk-space alert from any monitor — separate threshold noise from real pressure, read growth rate from history, rank consumer hypotheses.
High CPU/Memory Alert
Triage a CPU or memory threshold alert — separate a transient spike from sustained pressure via history, and route servers versus workstations differently.
Patch Failure Alert
Triage a patch-failure alert — separate a one-off from a repeat offender, detect reboot-pending as the usual culprit, correlate against the patch window.
RAID Degradation Alert
Triage a RAID degraded or failed-member alert with zero-margin urgency — one failure from data loss — and enforce the verify-backups-BEFORE-rebuild rule.
Was this page helpful?
⌘I