Continuous Improvement - Patterns
Learn more about Well-Architected Operational Excellence → Continuous Improvement
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | DORA | ✅ Deployment frequency is tracked; the fastest-moving teams deploy on demand, often multiple times per day |
| Platform | DORA | ✅ Lead time for changes is measured from code commit to production deployment; sub-day lead time signals high-performing delivery and sub-hour indicates exceptional capability |
| Platform | DORA | ✅ Change failure rate is tracked as the percentage of deployments causing production incidents requiring remediation, and kept consistently low |
| Platform | DORA | ✅ Failed deployment recovery time is measured; recovery under one hour signals delivery maturity |
| Platform | DORA | ✅ Deployment rework rate is tracked as the ratio of unplanned deployments resulting from production incidents |
| Platform | DORA | ✅ Metrics are organized into throughput (lead time, deployment frequency, failed deployment recovery time) and instability (change failure rate, rework rate) factors |
| Platform | Metrics | ✅ DORA metrics are measured continuously and trended over time, with improvement trajectory weighted more heavily than absolute values |
| Platform | Dashboards | ✅ Metrics are tracked in operational dashboards visible to the entire organization, creating transparency about operational performance |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | DORA | ⚠️ Low deployment frequency reflects deployment pain teams avoid, creating a vicious cycle where infrequent deployment makes each deployment riskier |
| Platform | DORA | ⚠️ Long lead times reflect excessive process overhead, insufficient automation, or organizational dysfunction |
| Platform | DORA | ⚠️ A high change failure rate reflects insufficient testing, inadequate deployment validation, or changes rushed without proper quality checks |
| Platform | DORA | ⚠️ Long recovery times reflect insufficient deployment validation, lack of automated rollback, or unclear ownership of production incidents |
| Platform | DORA | ⚠️ A high rework rate shows production incidents routinely driving emergency deployments, signaling gaps in pre-production testing, release validation, or change management |
| Platform | Metrics | ⚠️ Metrics are captured as point-in-time snapshots rather than trended over time, hiding whether operational performance is improving or degrading |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Process | ✅ Sprint retrospectives occur regularly (every sprint or monthly) creating a rhythm for reflection rather than waiting for crises |
| Platform | Process | ✅ Retrospectives use structured formats (Start/Stop/Continue, Mad/Sad/Glad, Timeline) that encourage participation and actionable outcomes |
| Platform | Process | ✅ Facilitators are rotated to prevent any single person from dominating discussion |
| Platform | Process | ✅ Retrospective insights are converted into concrete action items with owners and completion dates, producing a small number of improvements per session |
| Platform | Process | ✅ Action item completion is tracked across retrospectives, holding teams accountable to follow-through |
| Platform | Operations | ✅ Operational reviews assess aggregate operational health quarterly or monthly, examining trend data and comparing against objectives |
| Platform | Operations | ✅ Operational reviews engage leadership to secure resources for operational improvement work that competes against feature development |
| Platform | Operations | ✅ Reviews include operational metrics: availability against target by journey, latency trends by percentile, incident counts by severity with mean time to detect and resolve, deployment frequency and success rate, and operational load such as on-call pages and toil hours |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Process | ⚠️ Retrospectives wait for crises rather than occurring on a regular rhythm |
| Platform | Process | ⚠️ Retrospectives generate lengthy discussion but no action, wasting time and breeding cynicism |
| Platform | Process | ⚠️ Action items are defined without owners or completion dates, so improvements are discussed but never implemented |
| Platform | Process | ⚠️ Action item completion is never tracked across retrospectives, leaving follow-through unknown |
| Platform | Operations | ⚠️ Operations are treated as invisible background work receiving attention only during crises, with no regular operational review |
| Platform | Operations | ⚠️ Operational reviews lack leadership engagement, so improvement work loses to feature development when resources are allocated |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Operations | ✅ Operational Readiness Reviews assess whether new features or systems meet operational requirements before production launch |
| Platform | Operations | ✅ Reviews are conducted before production launch of new solutions, major features, or architectural changes with operational implications |
| Platform | Operations | ✅ Review timing is balanced so implementation is complete enough to assess without operational concerns arriving too late to address |
| Platform | Monitoring | ✅ The review confirms sufficient metrics are instrumented, alerts exist for failure scenarios, and dashboards show health status |
| Platform | Documentation | ✅ The review confirms runbooks exist for common operations, architecture is documented for responders, and escalation procedures are clear |
| Platform | Deployment | ✅ The review confirms deployment executes reliably, a rollback procedure exists, and deployment has been validated in staging |
| Platform | Performance | ✅ The review confirms performance targets are validated under realistic load, headroom exists for growth, and governor limit risks are assessed |
| Platform | Security | ✅ The review confirms security reviews are complete, audit requirements are met, and access control matches requirements |
| Platform | Dependencies | ✅ The review confirms external dependencies are identified, integration partners have service level agreements, and fallback behavior exists for dependency failures |
| Platform | Governance | ✅ Production launch is gated on review completion, making operational concerns deployment requirements rather than best-effort nice-to-haves |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Governance | ⚠️ New features launch to production without an operational readiness assessment, accumulating operational debt that makes future operational excellence increasingly difficult |
| Platform | Operations | ⚠️ The review is conducted too late, so operational concerns feel like deployment blockers and create pressure to skip fixes |
| Platform | Governance | ⚠️ Operational readiness is treated as a best-effort nice-to-have rather than a deployment requirement |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Culture | ✅ Psychological safety enables honest discussion of problems, mistakes, and near-misses without fear of punishment |
| Platform | Culture | ✅ Psychological safety is built through blameless postmortems, celebrating problem discovery, and leadership modeling vulnerability by discussing their own mistakes |
| Platform | Metrics | ✅ Measurement makes operational state visible through operational metrics, incident trends, DORA indicators, and user satisfaction scores, enabling objective prioritization of improvement work |
| Platform | Operations | ✅ Engineering capacity is deliberately reserved for operational improvement, technical debt reduction, and tooling and automation investment |
| Platform | Operations | ✅ Feedback loops connect operational experience to design decisions, so incidents drive architectural improvements, monitoring drives optimization, and deployment failures drive test coverage improvements |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Culture | ⚠️ Teams lacking psychological safety hide problems until they become catastrophic, preventing early intervention |
| Platform | Operations | ⚠️ Teams spend all capacity on features with no time for operational improvement, accumulating technical and operational debt that eventually forces crisis response |
| Platform | Operations | ⚠️ Operational experience is not fed back into design, so the same weaknesses recur and operations gradually degrade rather than improve |