Incident Management - Patterns
Learn more about Well-Architected Operational Excellence → Incident Management
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Monitoring | ✅ Alerting architecture includes severity classification routing critical incidents (production unavailable, data corruption, security compromise) to page on-call immediately |
| Platform | Monitoring | ✅ High severity incidents (degraded performance, partial outage) generate immediate notifications for prompt business-hours investigation |
| Platform | Monitoring | ✅ Warnings (approaching thresholds, configuration drift) aggregate into daily digest for proactive remediation |
| Platform | Monitoring | ✅ Alert enrichment includes diagnostic context: affected user count, error rate percentiles, recent deployments, related alerts, and direct dashboard links |
| Platform | Monitoring | ✅ Alert context reduces mean time to acknowledge by providing immediate situation assessment without gathering basic information through separate queries |
| Platform | Monitoring | ✅ Alert grouping prevents alert storms by consolidating hundreds of related alerts from single underlying problem into one notification with aggregate statistics |
| Platform | Monitoring | ✅ Grouped alerts show count, rate, and affected transactions revealing problem pattern without overwhelming responders |
| Platform | Monitoring | ✅ Alert suppression prevents known-transient conditions from generating alerts during expected maintenance windows or auto-remediation periods |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Monitoring | ⚠️ Integration failure generates 100s individual alerts in 5 minutes overwhelming on-call phone, each alert identical except timestamp |
| Platform | Monitoring | ⚠️ Alerts lack context about scope, recent changes, or problem patterns; on-call engineer must gather basic information manually |
| Platform | Monitoring | ⚠️ All alerts route through same channel regardless of severity; critical production outages compete with minor warnings |
| Platform | Monitoring | ⚠️ Alerts fire during scheduled maintenance windows generating false positives that erode trust in monitoring system |
| Platform | Monitoring | ⚠️ Alert grouping does not exist; each transaction failure fires separate alert creating alert storms during incidents |
| Platform | Monitoring | ⚠️ Alerts provide no links to relevant dashboards; responders spend time finding monitoring tools during incident response |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Operations | ✅ On-call rotation is weekly or bi-weekly balancing burden distribution against context preservation |
| Platform | Operations | ✅ Runbooks are documented for every alert type enabling on-call engineers to respond effectively without deep domain expertise |
| Platform | Operations | ✅ Runbooks include problem description, diagnostic steps, remediation procedures, escalation criteria, and related documentation links |
| Platform | Operations | ✅ Runbooks are living documents in version control that evolve through postmortem action items |
| Platform | Operations | ✅ Handoff procedures include brief meetings during rotation transitions transferring context including active issues, recent changes, known risks, and upcoming maintenance |
| Platform | Operations | ✅ Written handoff notes supplement verbal handoff preventing information loss during rotation transitions |
| Platform | Operations | ✅ Page frequency, after-hours incident count, and false positive rate are measured to identify unsustainable on-call burden requiring architectural remediation |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Operations | ⚠️ Runbooks do not exist; on-call engineers rely on tribal knowledge and frequent escalations to subject matter experts |
| Platform | Operations | ⚠️ Runbooks are stored in personal documents rather than shared version control where team members can access and update them |
| Platform | Operations | ⚠️ On-call rotation transitions happen without handoff meetings; incoming engineer lacks context about active issues and recent changes |
| Platform | Operations | ⚠️ High on-call burden continues indefinitely through operational heroics rather than driving architectural improvements |
| Platform | Operations | ⚠️ On-call metrics (page frequency, after-hours incidents, false positives) are not tracked; burden patterns remain invisible |
| Platform | Operations | ⚠️ Runbooks are created once during initial deployment and never updated as system evolves or new failure modes emerge |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Operations | ✅ Technical escalation paths define which subject matter experts handle specific problem domains |
| Platform | Operations | ✅ Management escalation paths define which managers receive notification for incidents exceeding severity or duration thresholds |
| Platform | Operations | ✅ Vendor escalation paths document Salesforce support case creation procedures and severity criteria |
| Platform | Operations | ✅ Executive escalation paths define which executives receive notification for business-critical incidents |
| Platform | Operations | ✅ Escalation criteria are explicit and objective: severity levels, time thresholds, affected user counts, business impact assessments |
| Platform | Operations | ✅ Escalation paths are documented in runbooks and accessible during incidents without searching for contact information |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Operations | ⚠️ Escalation paths are not documented; on-call engineers waste time searching for appropriate contacts during incidents |
| Platform | Operations | ⚠️ Escalation criteria are vague (“escalate if it seems bad”); responders lack clear guidance about when to escalate |
| Platform | Operations | ⚠️ Management notification happens too late because severity thresholds are undefined; executives learn about incidents from users |
| Platform | Operations | ⚠️ Salesforce support escalation procedures are not documented; teams delay opening support cases losing valuable response time |
| Platform | Operations | ⚠️ Subject matter expert contacts are outdated; escalations reach wrong people or fail during rotation changes |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Operations | ✅ Blameless post-incident reviews are conducted within 5 business days after significant operational events |
| Platform | Operations | ✅ Blameless culture focuses on system improvements rather than individual blame, creating psychological safety for honest assessment |
| Platform | Operations | ✅ Timeline includes detailed sequence from initial conditions through first signal, detection, acknowledgment, investigation, response, recovery, and validation |
| Platform | Operations | ✅ Timeline timestamps each event enabling duration analysis of detection, response, and recovery phases |
| Platform | Operations | ✅ Root cause identifies technical failure or condition that directly caused incident, distinguished from contributing factors |
| Platform | Operations | ✅ Contributing factors analysis reveals organizational, process, or architectural conditions that enabled or amplified incident impact |
| Platform | Operations | ✅ Detection analysis evaluates how incident was detected and whether detection could have been faster |
| Platform | Operations | ✅ Response evaluation identifies what went well during response and what was slower or harder than necessary |
| Platform | Operations | ✅ Action items are specific, assigned, and timebound preventing recurrence or improving future response |
| Platform | Operations | ✅ Action items avoid vague improvements (“improve monitoring”) in favor of concrete tasks (“add alert for API error rate exceeding 5% over 10min window”) |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Operations | ⚠️ Post-incident discussion concludes “batch job had a bug, we fixed it” with no documentation or systematic improvements |
| Platform | Operations | ⚠️ Post-incident reviews focus on individual fault rather than system improvements, destroying psychological safety |
| Platform | Operations | ⚠️ Timeline lacks timestamps preventing duration analysis; team cannot identify whether detection or response was slow |
| Platform | Operations | ⚠️ Root cause analysis does not distinguish immediate technical cause from organizational or process contributing factors |
| Platform | Operations | ⚠️ Detection analysis is skipped; monitoring gaps that allowed users to report incident first are not identified |
| Platform | Operations | ⚠️ Action items are vague without owners or due dates; improvements are discussed but never implemented |
| Platform | Operations | ⚠️ Post-incident reviews are not conducted within 5 business days; context is lost and urgency for improvements fades |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Knowledge Management | ✅ Post-incident reviews are documented and shared broadly beyond immediate incident responders for organizational learning |
| Platform | Knowledge Management | ✅ Post-incident reviews are stored in shared wiki or knowledge base with tagging enabling pattern analysis across multiple incidents |
| Platform | Knowledge Management | ✅ Action item completion is tracked in subsequent reviews preventing retrospectives from becoming repetitive discussion without follow-through |
| Platform | Knowledge Management | ✅ Action item completion rates are measured as operational health indicator revealing whether improvements are implemented or merely discussed |
| Platform | Operations | ✅ Incident patterns are analyzed across multiple incidents revealing systemic issues requiring architectural investment |
| Platform | Operations | ✅ Runbooks are updated based on post-incident learnings ensuring procedures reflect actual incident response experience |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Knowledge Management | ⚠️ Post-incident reviews are not documented; learnings remain with incident responders rather than spreading across organization |
| Platform | Knowledge Management | ⚠️ Post-incident reviews are stored in personal documents rather than shared knowledge base where team members can learn from history |
| Platform | Knowledge Management | ⚠️ Action items are defined but never tracked; completion rate is unknown and improvements are not implemented |
| Platform | Knowledge Management | ⚠️ Low action item completion rates persist without recognition that retrospectives are ineffective without follow-through |
| Platform | Operations | ⚠️ Incident patterns are not analyzed across multiple incidents; systemic issues remain hidden because each incident is treated in isolation |
| Platform | Operations | ⚠️ Runbooks are not updated after incidents; procedures become outdated as system evolves and new failure modes emerge |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Operations | ✅ Mean time to detect (MTTD) is tracked measuring duration from incident occurrence to detection |
| Platform | Operations | ✅ Mean time to acknowledge (MTTA) is tracked measuring duration from detection to on-call engineer acknowledgment |
| Platform | Operations | ✅ Mean time to restore (MTTR) is tracked measuring duration from detection to service restoration |
| Platform | Operations | ✅ Incident frequency is tracked with monthly trend review revealing whether stability is improving or degrading |
| Platform | Operations | ✅ Change failure rate is tracked measuring percentage of deployments causing production incidents requiring remediation |
| Platform | Operations | ✅ Change failure rate is kept consistently low, a small fraction of all deployments |
| Platform | Operations | ✅ Incident metrics are shared with development teams during sprint planning informing backlog prioritization and architectural improvements |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Operations | ⚠️ Incident metrics are not tracked; team has no visibility into whether response times are improving or degrading |
| Platform | Operations | ⚠️ MTTD, MTTA, and MTTR are not measured; teams cannot identify whether detection or response is the bottleneck |
| Platform | Operations | ⚠️ Change failure rate is not tracked; team does not know whether quality gates are effective |
| Platform | Operations | ⚠️ Incident frequency trends are not reviewed monthly; stability degradation is not detected until major incident occurs |
| Platform | Operations | ⚠️ Incident metrics remain in operations team without sharing to development teams; architectural problems are not addressed |
| Platform | Operations | ⚠️ Metrics are tracked but never used to inform decisions; measurement without action wastes effort and provides no value |