Observability and Monitoring - Patterns
Learn more about Well-Architected Operational Excellence → Observability and Monitoring
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Event Monitoring | ✅ Event Monitoring is designed into initial architecture capturing correlation IDs that trace requests through async processing chains and external integrations |
| Platform | Event Monitoring | ✅ Observability infrastructure is included from initial sprint rather than retrofitted after production incidents prove impossible to diagnose |
| Platform | Integration | ✅ Integration checkpoints are placed deliberately to enable request tracing across system boundaries |
| Platform | Platform Events | ✅ Platform Event payloads are structured for operational visibility with transaction markers and context |
| Platform | Custom Logging | ✅ Custom logging framework captures discrete events with full contextual information including user identity, timestamp, duration, and outcome |
| Platform | Metrics | ✅ Metrics aggregate numerical measurements over time revealing trends including API consumption rates, Apex CPU time distributions, and batch job success rates |
| Platform | Distributed Tracing | ✅ Traces connect synchronous API calls to asynchronous processing chains, platform events to subscriber executions, and integration requests to external system responses |
| Platform | Observability | ✅ Observability enables operators to answer arbitrary questions about system behavior without deploying new instrumentation for each investigation |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Event Monitoring | ⚠️ Solution launches without custom logging, requiring major refactoring six months later when production incidents prove impossible to diagnose with platform logs alone |
| Platform | Event Monitoring | ⚠️ Observability is retrofitted into existing solution requiring instrumentation changes touching most components and risking bugs during operational improvement work |
| Platform | Integration | ⚠️ Integration failures occur but request paths cannot be traced across system boundaries without comprehensive telemetry |
| Platform | Platform Events | ⚠️ Platform Event payloads lack operational context making incident investigation dependent on searching multiple systems |
| Platform | Custom Logging | ⚠️ Logs capture insufficient context; investigations require reproducing issues to gather additional information |
| Platform | Observability | ⚠️ Predefined dashboards cannot answer unexpected questions about system behavior emerging in production |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Event Monitoring | ✅ Event Monitoring is enabled for production environments capturing API usage, login activity, Apex execution, SOQL queries, Visualforce pages, Lightning pages, and report runs |
| Platform | Event Monitoring | ✅ Event Log Files export hourly to centralized log platform enabling historical analysis beyond native platform retention |
| Platform | Event Monitoring | ✅ External log aggregation enables correlation with enterprise telemetry from other systems for end-to-end transaction tracing |
| Platform | Event Monitoring | ✅ Event Log Files are retained matching regulatory requirements for compliance reporting and long-term trend analysis |
| Platform | Proactive Monitoring | ✅ Proactive Monitoring thresholds are tuned deliberately based on org-specific normal operating range rather than accepting defaults |
| Platform | Proactive Monitoring | ✅ Proactive Monitoring alerts detect API limit spikes, concurrent Apex failures, SOQL row limit approaches, and row lock contention before user-visible incidents |
| Platform | Proactive Monitoring | ✅ Tuned thresholds generate actionable signals that teams trust and act on creating virtuous cycles of continuous improvement |
| Platform | Scale Center | ✅ Scale Center baselines are established during solution stabilization and revisited after each major release |
| Platform | Scale Center | ✅ Scale Center reveals which operations consume the most resources, which transactions approach timeout thresholds, and where optimization investment yields greatest impact |
| Platform | Scale Center | ✅ Transaction time increases from baseline prompt investigation revealing unintentional N+1 query patterns introduced in latest deployment |
| Platform | Setup Audit Trail | ✅ Setup Audit Trail tracks configuration changes including permission modifications, metadata deployments, and security setting updates with retention up to 180 days |
| Platform | Setup Audit Trail | ✅ Setup Audit Trail entries are exported for retention beyond 180 days when compliance or contractual requirements demand longer historical windows |
| Platform | Shield | ✅ Field Audit Trail is enabled selectively for fields containing sensitive data, regulated data requiring change history, or critical business data |
| Platform | Shield | ✅ Field Audit Trail provides up to 10 years retention for fields where understanding historical values aids operations and compliance reporting |
| Platform | Health Check | ✅ Health Check reviews are scheduled quarterly with findings remediated based on risk prioritization for the environment |
| Platform | Health Check | ✅ Health Check findings receive deliberate review rather than passive acceptance; compensating controls or different risk tolerances are documented |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Event Monitoring | ⚠️ Event Monitoring is enabled but logs are never exported, limiting investigations to the native retention window and preventing historical trend analysis |
| Platform | Event Monitoring | ⚠️ Event Log Files remain in platform storage; external aggregation is not implemented preventing long-term capacity planning |
| Platform | Proactive Monitoring | ⚠️ Proactive Monitoring uses default thresholds that are too sensitive or too lax, generating noisy alerts that teams learn to ignore |
| Platform | Proactive Monitoring | ⚠️ Alert noise destroys trust in monitoring system; teams stop responding to valid signals |
| Platform | Scale Center | ⚠️ Scale Center is enabled but never reviewed proactively; performance slowly degrades until users complain after months of incremental degradation |
| Platform | Scale Center | ⚠️ Performance metrics lack baseline context; unclear whether 3-second transaction times are typical, improving, or degrading |
| Platform | Setup Audit Trail | ⚠️ Setup Audit Trail is not exported; retention limitations prevent security investigations or compliance validation beyond 180 days |
| Platform | Shield | ⚠️ Field Audit Trail is enabled for all fields rather than selectively, creating storage costs without corresponding compliance value |
| Platform | Health Check | ⚠️ Health Check findings are never reviewed; security configuration drifts from baseline recommendations without deliberate acceptance |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | User Flows | ✅ Critical user flows are defined based on business impact and monitored end-to-end for success rates, completion time, abandonment points, and error rates |
| Platform | User Flows | ✅ Critical flows include revenue-generating activities (order submission, contract execution, opportunity close), high-volume activities (login, search, record creation), and compliance-required activities (consent capture, data subject rights fulfillment) |
| Platform | User Flows | ✅ Flows are instrumented with transaction markers indicating start, completion, abandonment, and failure at each significant step |
| Platform | User Flows | ✅ Order submission flow instrumented with platform events at each stage (cart validation, inventory check, payment processing, fulfillment trigger) showing success rate by stage |
| Platform | User Flows | ✅ Flow-level monitoring reveals problems invisible in component-level monitoring because failures can occur at any transition point |
| Platform | User Flows | ✅ Alerts fire when flow success rates drop below acceptable thresholds or duration exceeds latency targets |
| Platform | Integration | ✅ Integration health monitoring detects external system failures, network issues, authentication problems, and rate limit approaches before cascading into user-facing failures |
| Platform | Integration | ✅ Success rate monitoring alerts when falling below 99% for critical synchronous integrations |
| Platform | Integration | ✅ Error distribution monitoring alerts when new error types appear that did not exist in baseline |
| Platform | Integration | ✅ Latency percentile monitoring (p95) alerts when exceeding 3 seconds for synchronous user-facing calls |
| Platform | Integration | ✅ Throughput monitoring alerts when above 80% of known rate limit or SLA threshold |
| Platform | Integration | ✅ Circuit breaker state monitoring alerts when any circuit opens indicating external system degradation |
| Platform | Integration | ✅ Retry rate monitoring alerts when sustained retry rate exceeds 5% indicating reliability issues |
| Platform | Integration | ✅ Integration telemetry routes to centralized monitoring platforms enabling correlation across system boundaries |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | User Flows | ⚠️ Components are monitored independently showing healthy operation while 12% of users abandon critical flows with no visibility into why |
| Platform | User Flows | ⚠️ Component health dashboards show green while users cannot complete critical workflows revealing monitoring architecture failure |
| Platform | User Flows | ⚠️ User journey touching multiple Apex classes, flows, platform events, and external integrations lacks end-to-end visibility |
| Platform | User Flows | ⚠️ Critical business journeys lack instrumentation; success rates and abandonment points are unknown |
| Platform | Integration | ⚠️ Integration failures are detected only after cascading into user-facing failures rather than proactively |
| Platform | Integration | ⚠️ Integration telemetry remains isolated in individual systems preventing correlation and root cause analysis across boundaries |
| Platform | Integration | ⚠️ Monitoring alerts on downstream symptoms after damage is done rather than revealing root causes early |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | SLIs | ✅ Availability SLIs measure percentage of time critical flows execute successfully from user perspective (receiving expected result, not just 200 status code) |
| Platform | SLIs | ✅ Latency SLIs measure time to complete operations at p50 (median), p90 (90th percentile excluding outliers), p95 (95th percentile excluding worst outliers), and p99 (99th percentile showing worst-case experience) |
| Platform | SLIs | ✅ Latency targets vary by operation criticality and user expectation (500ms search feels instantaneous, same latency for typing input feels sluggish) |
| Platform | SLIs | ✅ Throughput SLIs measure volume of successful operations per time unit revealing capacity utilization and growth trends |
| Platform | SLIs | ✅ Throughput combined with latency reveals whether system maintains performance under increasing load or degrades as volume grows |
| Platform | SLIs | ✅ Error rate SLIs distinguish between error types (user errors like validation failures versus system errors like governor limit exceptions) |
| Platform | SLOs | ✅ Service Level Objectives are established defining acceptable SLI thresholds based on business requirements before selecting architecture patterns |
| Platform | SLOs | ✅ Business requirement “Sales reps must see opportunity updates within 5 seconds” drives SLO “Opportunity detail page loads complete within 3 seconds at p95” which drives architecture decisions |
| Platform | SLOs | ✅ SLOs emerge from business requirements and drive architectural decisions rather than being defined based on what system currently achieves |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | SLIs | ⚠️ Availability is measured from technical perspective (200 status code) rather than user perspective (receiving expected result) |
| Platform | SLIs | ⚠️ Latency is measured only at median without understanding worst-case user experience at p95 or p99 percentiles |
| Platform | SLIs | ⚠️ Error rate SLIs do not distinguish between user errors and system errors, masking different problems requiring different solutions |
| Platform | SLOs | ⚠️ Architecture is built first, then SLO is defined as “whatever we currently achieve” leading to acceptance of 12-second page loads because system design never considered user requirements |
| Platform | SLOs | ⚠️ SLOs are architecture-driven rather than business-driven, setting targets based on what system achieves rather than what users need |
Patterns
| Where to look | What good looks like |
|---|---|
| Platform | Alerting | ✅ Severity classification routes critical incidents (production unavailable, data corruption risk, security compromise) to page on-call engineers immediately |
| Platform | Alerting | ✅ High severity incidents (degraded performance, partial outages) generate notifications for prompt investigation within business hours |
| Platform | Alerting | ✅ Warnings (approaching thresholds, configuration drift) generate daily digest notifications for proactive remediation |
| Platform | Alerting | ✅ Alert context includes affected user count, error messages, stack traces, recent deployments, related alerts, and links to relevant dashboards |
| Platform | Alerting | ✅ Rich context reduces mean time to acknowledge by enabling immediate severity assessment and investigation without gathering basic information |
| Platform | Alerting | ✅ Alert grouping prevents alert storms where hundreds of related alerts fire for single underlying problem |
| Platform | Alerting | ✅ Integration failure triggers one alert grouped by integration endpoint showing 156 failed requests in past 5 minutes with latency graph, error distribution, and dashboard link |
| Platform | Alerting | ✅ Alert suppression prevents known-transient conditions from generating alerts during expected maintenance windows or auto-remediation periods |
| Platform | Alerting | ✅ Alerts are suppressed during scheduled maintenance, during first few minutes after deployment while caches warm, and during brief transient failures where retry logic succeeds |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Platform | Alerting | ⚠️ All alerts route to same channel without severity classification; critical incidents are not distinguished from routine warnings |
| Platform | Alerting | ⚠️ Alerts lack diagnostic context; on-call engineers must gather basic information through separate queries before assessing severity |
| Platform | Alerting | ⚠️ Integration failure generates 156 individual alerts in 5 minutes overwhelming on-call engineer’s phone, each alert identical except for timestamp |
| Platform | Alerting | ⚠️ Alert storms prevent responders from distinguishing single underlying problem from hundreds of related symptoms |
| Platform | Alerting | ⚠️ Alerts fire during known maintenance windows and transient conditions generating noise that erodes trust in monitoring |
| Platform | Alerting | ⚠️ Alert fatigue from noisy false positives causes teams to ignore notifications reducing response effectiveness |