Incident Management - Patterns

Learn more about Well-Architected Operational Excellence → Incident Management

Patterns

Where to lookWhat good looks like
Platform | Monitoring✅ Alerting architecture includes severity classification routing critical incidents (production unavailable, data corruption, security compromise) to page on-call immediately
Platform | Monitoring✅ High severity incidents (degraded performance, partial outage) generate immediate notifications for prompt business-hours investigation
Platform | Monitoring✅ Warnings (approaching thresholds, configuration drift) aggregate into daily digest for proactive remediation
Platform | Monitoring✅ Alert enrichment includes diagnostic context: affected user count, error rate percentiles, recent deployments, related alerts, and direct dashboard links
Platform | Monitoring✅ Alert context reduces mean time to acknowledge by providing immediate situation assessment without gathering basic information through separate queries
Platform | Monitoring✅ Alert grouping prevents alert storms by consolidating hundreds of related alerts from single underlying problem into one notification with aggregate statistics
Platform | Monitoring✅ Grouped alerts show count, rate, and affected transactions revealing problem pattern without overwhelming responders
Platform | Monitoring✅ Alert suppression prevents known-transient conditions from generating alerts during expected maintenance windows or auto-remediation periods

Anti-Patterns

Where to lookWhat bad looks like
Platform | Monitoring⚠️ Integration failure generates 100s individual alerts in 5 minutes overwhelming on-call phone, each alert identical except timestamp
Platform | Monitoring⚠️ Alerts lack context about scope, recent changes, or problem patterns; on-call engineer must gather basic information manually
Platform | Monitoring⚠️ All alerts route through same channel regardless of severity; critical production outages compete with minor warnings
Platform | Monitoring⚠️ Alerts fire during scheduled maintenance windows generating false positives that erode trust in monitoring system
Platform | Monitoring⚠️ Alert grouping does not exist; each transaction failure fires separate alert creating alert storms during incidents
Platform | Monitoring⚠️ Alerts provide no links to relevant dashboards; responders spend time finding monitoring tools during incident response

Patterns

Where to lookWhat good looks like
Platform | Operations✅ On-call rotation is weekly or bi-weekly balancing burden distribution against context preservation
Platform | Operations✅ Runbooks are documented for every alert type enabling on-call engineers to respond effectively without deep domain expertise
Platform | Operations✅ Runbooks include problem description, diagnostic steps, remediation procedures, escalation criteria, and related documentation links
Platform | Operations✅ Runbooks are living documents in version control that evolve through postmortem action items
Platform | Operations✅ Handoff procedures include brief meetings during rotation transitions transferring context including active issues, recent changes, known risks, and upcoming maintenance
Platform | Operations✅ Written handoff notes supplement verbal handoff preventing information loss during rotation transitions
Platform | Operations✅ Page frequency, after-hours incident count, and false positive rate are measured to identify unsustainable on-call burden requiring architectural remediation

Anti-Patterns

Where to lookWhat bad looks like
Platform | Operations⚠️ Runbooks do not exist; on-call engineers rely on tribal knowledge and frequent escalations to subject matter experts
Platform | Operations⚠️ Runbooks are stored in personal documents rather than shared version control where team members can access and update them
Platform | Operations⚠️ On-call rotation transitions happen without handoff meetings; incoming engineer lacks context about active issues and recent changes
Platform | Operations⚠️ High on-call burden continues indefinitely through operational heroics rather than driving architectural improvements
Platform | Operations⚠️ On-call metrics (page frequency, after-hours incidents, false positives) are not tracked; burden patterns remain invisible
Platform | Operations⚠️ Runbooks are created once during initial deployment and never updated as system evolves or new failure modes emerge

Patterns

Where to lookWhat good looks like
Platform | Operations✅ Technical escalation paths define which subject matter experts handle specific problem domains
Platform | Operations✅ Management escalation paths define which managers receive notification for incidents exceeding severity or duration thresholds
Platform | Operations✅ Vendor escalation paths document Salesforce support case creation procedures and severity criteria
Platform | Operations✅ Executive escalation paths define which executives receive notification for business-critical incidents
Platform | Operations✅ Escalation criteria are explicit and objective: severity levels, time thresholds, affected user counts, business impact assessments
Platform | Operations✅ Escalation paths are documented in runbooks and accessible during incidents without searching for contact information

Anti-Patterns

Where to lookWhat bad looks like
Platform | Operations⚠️ Escalation paths are not documented; on-call engineers waste time searching for appropriate contacts during incidents
Platform | Operations⚠️ Escalation criteria are vague (“escalate if it seems bad”); responders lack clear guidance about when to escalate
Platform | Operations⚠️ Management notification happens too late because severity thresholds are undefined; executives learn about incidents from users
Platform | Operations⚠️ Salesforce support escalation procedures are not documented; teams delay opening support cases losing valuable response time
Platform | Operations⚠️ Subject matter expert contacts are outdated; escalations reach wrong people or fail during rotation changes

Patterns

Where to lookWhat good looks like
Platform | Operations✅ Blameless post-incident reviews are conducted within 5 business days after significant operational events
Platform | Operations✅ Blameless culture focuses on system improvements rather than individual blame, creating psychological safety for honest assessment
Platform | Operations✅ Timeline includes detailed sequence from initial conditions through first signal, detection, acknowledgment, investigation, response, recovery, and validation
Platform | Operations✅ Timeline timestamps each event enabling duration analysis of detection, response, and recovery phases
Platform | Operations✅ Root cause identifies technical failure or condition that directly caused incident, distinguished from contributing factors
Platform | Operations✅ Contributing factors analysis reveals organizational, process, or architectural conditions that enabled or amplified incident impact
Platform | Operations✅ Detection analysis evaluates how incident was detected and whether detection could have been faster
Platform | Operations✅ Response evaluation identifies what went well during response and what was slower or harder than necessary
Platform | Operations✅ Action items are specific, assigned, and timebound preventing recurrence or improving future response
Platform | Operations✅ Action items avoid vague improvements (“improve monitoring”) in favor of concrete tasks (“add alert for API error rate exceeding 5% over 10min window”)

Anti-Patterns

Where to lookWhat bad looks like
Platform | Operations⚠️ Post-incident discussion concludes “batch job had a bug, we fixed it” with no documentation or systematic improvements
Platform | Operations⚠️ Post-incident reviews focus on individual fault rather than system improvements, destroying psychological safety
Platform | Operations⚠️ Timeline lacks timestamps preventing duration analysis; team cannot identify whether detection or response was slow
Platform | Operations⚠️ Root cause analysis does not distinguish immediate technical cause from organizational or process contributing factors
Platform | Operations⚠️ Detection analysis is skipped; monitoring gaps that allowed users to report incident first are not identified
Platform | Operations⚠️ Action items are vague without owners or due dates; improvements are discussed but never implemented
Platform | Operations⚠️ Post-incident reviews are not conducted within 5 business days; context is lost and urgency for improvements fades

Patterns

Where to lookWhat good looks like
Platform | Knowledge Management✅ Post-incident reviews are documented and shared broadly beyond immediate incident responders for organizational learning
Platform | Knowledge Management✅ Post-incident reviews are stored in shared wiki or knowledge base with tagging enabling pattern analysis across multiple incidents
Platform | Knowledge Management✅ Action item completion is tracked in subsequent reviews preventing retrospectives from becoming repetitive discussion without follow-through
Platform | Knowledge Management✅ Action item completion rates are measured as operational health indicator revealing whether improvements are implemented or merely discussed
Platform | Operations✅ Incident patterns are analyzed across multiple incidents revealing systemic issues requiring architectural investment
Platform | Operations✅ Runbooks are updated based on post-incident learnings ensuring procedures reflect actual incident response experience

Anti-Patterns

Where to lookWhat bad looks like
Platform | Knowledge Management⚠️ Post-incident reviews are not documented; learnings remain with incident responders rather than spreading across organization
Platform | Knowledge Management⚠️ Post-incident reviews are stored in personal documents rather than shared knowledge base where team members can learn from history
Platform | Knowledge Management⚠️ Action items are defined but never tracked; completion rate is unknown and improvements are not implemented
Platform | Knowledge Management⚠️ Low action item completion rates persist without recognition that retrospectives are ineffective without follow-through
Platform | Operations⚠️ Incident patterns are not analyzed across multiple incidents; systemic issues remain hidden because each incident is treated in isolation
Platform | Operations⚠️ Runbooks are not updated after incidents; procedures become outdated as system evolves and new failure modes emerge

Patterns

Where to lookWhat good looks like
Platform | Operations✅ Mean time to detect (MTTD) is tracked measuring duration from incident occurrence to detection
Platform | Operations✅ Mean time to acknowledge (MTTA) is tracked measuring duration from detection to on-call engineer acknowledgment
Platform | Operations✅ Mean time to restore (MTTR) is tracked measuring duration from detection to service restoration
Platform | Operations✅ Incident frequency is tracked with monthly trend review revealing whether stability is improving or degrading
Platform | Operations✅ Change failure rate is tracked measuring percentage of deployments causing production incidents requiring remediation
Platform | Operations✅ Change failure rate is kept consistently low, a small fraction of all deployments
Platform | Operations✅ Incident metrics are shared with development teams during sprint planning informing backlog prioritization and architectural improvements

Anti-Patterns

Where to lookWhat bad looks like
Platform | Operations⚠️ Incident metrics are not tracked; team has no visibility into whether response times are improving or degrading
Platform | Operations⚠️ MTTD, MTTA, and MTTR are not measured; teams cannot identify whether detection or response is the bottleneck
Platform | Operations⚠️ Change failure rate is not tracked; team does not know whether quality gates are effective
Platform | Operations⚠️ Incident frequency trends are not reviewed monthly; stability degradation is not detected until major incident occurs
Platform | Operations⚠️ Incident metrics remain in operations team without sharing to development teams; architectural problems are not addressed
Platform | Operations⚠️ Metrics are tracked but never used to inform decisions; measurement without action wastes effort and provides no value