Continuous Improvement - Patterns

Learn more about Well-Architected Operational Excellence → Continuous Improvement

Patterns

Where to lookWhat good looks like
Platform | DORA✅ Deployment frequency is tracked; the fastest-moving teams deploy on demand, often multiple times per day
Platform | DORA✅ Lead time for changes is measured from code commit to production deployment; sub-day lead time signals high-performing delivery and sub-hour indicates exceptional capability
Platform | DORA✅ Change failure rate is tracked as the percentage of deployments causing production incidents requiring remediation, and kept consistently low
Platform | DORA✅ Failed deployment recovery time is measured; recovery under one hour signals delivery maturity
Platform | DORA✅ Deployment rework rate is tracked as the ratio of unplanned deployments resulting from production incidents
Platform | DORA✅ Metrics are organized into throughput (lead time, deployment frequency, failed deployment recovery time) and instability (change failure rate, rework rate) factors
Platform | Metrics✅ DORA metrics are measured continuously and trended over time, with improvement trajectory weighted more heavily than absolute values
Platform | Dashboards✅ Metrics are tracked in operational dashboards visible to the entire organization, creating transparency about operational performance

Anti-Patterns

Where to lookWhat bad looks like
Platform | DORA⚠️ Low deployment frequency reflects deployment pain teams avoid, creating a vicious cycle where infrequent deployment makes each deployment riskier
Platform | DORA⚠️ Long lead times reflect excessive process overhead, insufficient automation, or organizational dysfunction
Platform | DORA⚠️ A high change failure rate reflects insufficient testing, inadequate deployment validation, or changes rushed without proper quality checks
Platform | DORA⚠️ Long recovery times reflect insufficient deployment validation, lack of automated rollback, or unclear ownership of production incidents
Platform | DORA⚠️ A high rework rate shows production incidents routinely driving emergency deployments, signaling gaps in pre-production testing, release validation, or change management
Platform | Metrics⚠️ Metrics are captured as point-in-time snapshots rather than trended over time, hiding whether operational performance is improving or degrading

Patterns

Where to lookWhat good looks like
Platform | Process✅ Sprint retrospectives occur regularly (every sprint or monthly) creating a rhythm for reflection rather than waiting for crises
Platform | Process✅ Retrospectives use structured formats (Start/Stop/Continue, Mad/Sad/Glad, Timeline) that encourage participation and actionable outcomes
Platform | Process✅ Facilitators are rotated to prevent any single person from dominating discussion
Platform | Process✅ Retrospective insights are converted into concrete action items with owners and completion dates, producing a small number of improvements per session
Platform | Process✅ Action item completion is tracked across retrospectives, holding teams accountable to follow-through
Platform | Operations✅ Operational reviews assess aggregate operational health quarterly or monthly, examining trend data and comparing against objectives
Platform | Operations✅ Operational reviews engage leadership to secure resources for operational improvement work that competes against feature development
Platform | Operations✅ Reviews include operational metrics: availability against target by journey, latency trends by percentile, incident counts by severity with mean time to detect and resolve, deployment frequency and success rate, and operational load such as on-call pages and toil hours

Anti-Patterns

Where to lookWhat bad looks like
Platform | Process⚠️ Retrospectives wait for crises rather than occurring on a regular rhythm
Platform | Process⚠️ Retrospectives generate lengthy discussion but no action, wasting time and breeding cynicism
Platform | Process⚠️ Action items are defined without owners or completion dates, so improvements are discussed but never implemented
Platform | Process⚠️ Action item completion is never tracked across retrospectives, leaving follow-through unknown
Platform | Operations⚠️ Operations are treated as invisible background work receiving attention only during crises, with no regular operational review
Platform | Operations⚠️ Operational reviews lack leadership engagement, so improvement work loses to feature development when resources are allocated

Patterns

Where to lookWhat good looks like
Platform | Operations✅ Operational Readiness Reviews assess whether new features or systems meet operational requirements before production launch
Platform | Operations✅ Reviews are conducted before production launch of new solutions, major features, or architectural changes with operational implications
Platform | Operations✅ Review timing is balanced so implementation is complete enough to assess without operational concerns arriving too late to address
Platform | Monitoring✅ The review confirms sufficient metrics are instrumented, alerts exist for failure scenarios, and dashboards show health status
Platform | Documentation✅ The review confirms runbooks exist for common operations, architecture is documented for responders, and escalation procedures are clear
Platform | Deployment✅ The review confirms deployment executes reliably, a rollback procedure exists, and deployment has been validated in staging
Platform | Performance✅ The review confirms performance targets are validated under realistic load, headroom exists for growth, and governor limit risks are assessed
Platform | Security✅ The review confirms security reviews are complete, audit requirements are met, and access control matches requirements
Platform | Dependencies✅ The review confirms external dependencies are identified, integration partners have service level agreements, and fallback behavior exists for dependency failures
Platform | Governance✅ Production launch is gated on review completion, making operational concerns deployment requirements rather than best-effort nice-to-haves

Anti-Patterns

Where to lookWhat bad looks like
Platform | Governance⚠️ New features launch to production without an operational readiness assessment, accumulating operational debt that makes future operational excellence increasingly difficult
Platform | Operations⚠️ The review is conducted too late, so operational concerns feel like deployment blockers and create pressure to skip fixes
Platform | Governance⚠️ Operational readiness is treated as a best-effort nice-to-have rather than a deployment requirement

Patterns

Where to lookWhat good looks like
Platform | Culture✅ Psychological safety enables honest discussion of problems, mistakes, and near-misses without fear of punishment
Platform | Culture✅ Psychological safety is built through blameless postmortems, celebrating problem discovery, and leadership modeling vulnerability by discussing their own mistakes
Platform | Metrics✅ Measurement makes operational state visible through operational metrics, incident trends, DORA indicators, and user satisfaction scores, enabling objective prioritization of improvement work
Platform | Operations✅ Engineering capacity is deliberately reserved for operational improvement, technical debt reduction, and tooling and automation investment
Platform | Operations✅ Feedback loops connect operational experience to design decisions, so incidents drive architectural improvements, monitoring drives optimization, and deployment failures drive test coverage improvements

Anti-Patterns

Where to lookWhat bad looks like
Platform | Culture⚠️ Teams lacking psychological safety hide problems until they become catastrophic, preventing early intervention
Platform | Operations⚠️ Teams spend all capacity on features with no time for operational improvement, accumulating technical and operational debt that eventually forces crisis response
Platform | Operations⚠️ Operational experience is not fed back into design, so the same weaknesses recur and operations gradually degrade rather than improve