Observability and Monitoring - Patterns

Learn more about Well-Architected Operational Excellence → Observability and Monitoring

Patterns

Where to lookWhat good looks like
Platform | Event Monitoring✅ Event Monitoring is designed into initial architecture capturing correlation IDs that trace requests through async processing chains and external integrations
Platform | Event Monitoring✅ Observability infrastructure is included from initial sprint rather than retrofitted after production incidents prove impossible to diagnose
Platform | Integration✅ Integration checkpoints are placed deliberately to enable request tracing across system boundaries
Platform | Platform Events✅ Platform Event payloads are structured for operational visibility with transaction markers and context
Platform | Custom Logging✅ Custom logging framework captures discrete events with full contextual information including user identity, timestamp, duration, and outcome
Platform | Metrics✅ Metrics aggregate numerical measurements over time revealing trends including API consumption rates, Apex CPU time distributions, and batch job success rates
Platform | Distributed Tracing✅ Traces connect synchronous API calls to asynchronous processing chains, platform events to subscriber executions, and integration requests to external system responses
Platform | Observability✅ Observability enables operators to answer arbitrary questions about system behavior without deploying new instrumentation for each investigation

Anti-Patterns

Where to lookWhat bad looks like
Platform | Event Monitoring⚠️ Solution launches without custom logging, requiring major refactoring six months later when production incidents prove impossible to diagnose with platform logs alone
Platform | Event Monitoring⚠️ Observability is retrofitted into existing solution requiring instrumentation changes touching most components and risking bugs during operational improvement work
Platform | Integration⚠️ Integration failures occur but request paths cannot be traced across system boundaries without comprehensive telemetry
Platform | Platform Events⚠️ Platform Event payloads lack operational context making incident investigation dependent on searching multiple systems
Platform | Custom Logging⚠️ Logs capture insufficient context; investigations require reproducing issues to gather additional information
Platform | Observability⚠️ Predefined dashboards cannot answer unexpected questions about system behavior emerging in production

Patterns

Where to lookWhat good looks like
Platform | Event Monitoring✅ Event Monitoring is enabled for production environments capturing API usage, login activity, Apex execution, SOQL queries, Visualforce pages, Lightning pages, and report runs
Platform | Event Monitoring✅ Event Log Files export hourly to centralized log platform enabling historical analysis beyond native platform retention
Platform | Event Monitoring✅ External log aggregation enables correlation with enterprise telemetry from other systems for end-to-end transaction tracing
Platform | Event Monitoring✅ Event Log Files are retained matching regulatory requirements for compliance reporting and long-term trend analysis
Platform | Proactive Monitoring✅ Proactive Monitoring thresholds are tuned deliberately based on org-specific normal operating range rather than accepting defaults
Platform | Proactive Monitoring✅ Proactive Monitoring alerts detect API limit spikes, concurrent Apex failures, SOQL row limit approaches, and row lock contention before user-visible incidents
Platform | Proactive Monitoring✅ Tuned thresholds generate actionable signals that teams trust and act on creating virtuous cycles of continuous improvement
Platform | Scale Center✅ Scale Center baselines are established during solution stabilization and revisited after each major release
Platform | Scale Center✅ Scale Center reveals which operations consume the most resources, which transactions approach timeout thresholds, and where optimization investment yields greatest impact
Platform | Scale Center✅ Transaction time increases from baseline prompt investigation revealing unintentional N+1 query patterns introduced in latest deployment
Platform | Setup Audit Trail✅ Setup Audit Trail tracks configuration changes including permission modifications, metadata deployments, and security setting updates with retention up to 180 days
Platform | Setup Audit Trail✅ Setup Audit Trail entries are exported for retention beyond 180 days when compliance or contractual requirements demand longer historical windows
Platform | Shield✅ Field Audit Trail is enabled selectively for fields containing sensitive data, regulated data requiring change history, or critical business data
Platform | Shield✅ Field Audit Trail provides up to 10 years retention for fields where understanding historical values aids operations and compliance reporting
Platform | Health Check✅ Health Check reviews are scheduled quarterly with findings remediated based on risk prioritization for the environment
Platform | Health Check✅ Health Check findings receive deliberate review rather than passive acceptance; compensating controls or different risk tolerances are documented

Anti-Patterns

Where to lookWhat bad looks like
Platform | Event Monitoring⚠️ Event Monitoring is enabled but logs are never exported, limiting investigations to the native retention window and preventing historical trend analysis
Platform | Event Monitoring⚠️ Event Log Files remain in platform storage; external aggregation is not implemented preventing long-term capacity planning
Platform | Proactive Monitoring⚠️ Proactive Monitoring uses default thresholds that are too sensitive or too lax, generating noisy alerts that teams learn to ignore
Platform | Proactive Monitoring⚠️ Alert noise destroys trust in monitoring system; teams stop responding to valid signals
Platform | Scale Center⚠️ Scale Center is enabled but never reviewed proactively; performance slowly degrades until users complain after months of incremental degradation
Platform | Scale Center⚠️ Performance metrics lack baseline context; unclear whether 3-second transaction times are typical, improving, or degrading
Platform | Setup Audit Trail⚠️ Setup Audit Trail is not exported; retention limitations prevent security investigations or compliance validation beyond 180 days
Platform | Shield⚠️ Field Audit Trail is enabled for all fields rather than selectively, creating storage costs without corresponding compliance value
Platform | Health Check⚠️ Health Check findings are never reviewed; security configuration drifts from baseline recommendations without deliberate acceptance

Patterns

Where to lookWhat good looks like
Platform | User Flows✅ Critical user flows are defined based on business impact and monitored end-to-end for success rates, completion time, abandonment points, and error rates
Platform | User Flows✅ Critical flows include revenue-generating activities (order submission, contract execution, opportunity close), high-volume activities (login, search, record creation), and compliance-required activities (consent capture, data subject rights fulfillment)
Platform | User Flows✅ Flows are instrumented with transaction markers indicating start, completion, abandonment, and failure at each significant step
Platform | User Flows✅ Order submission flow instrumented with platform events at each stage (cart validation, inventory check, payment processing, fulfillment trigger) showing success rate by stage
Platform | User Flows✅ Flow-level monitoring reveals problems invisible in component-level monitoring because failures can occur at any transition point
Platform | User Flows✅ Alerts fire when flow success rates drop below acceptable thresholds or duration exceeds latency targets
Platform | Integration✅ Integration health monitoring detects external system failures, network issues, authentication problems, and rate limit approaches before cascading into user-facing failures
Platform | Integration✅ Success rate monitoring alerts when falling below 99% for critical synchronous integrations
Platform | Integration✅ Error distribution monitoring alerts when new error types appear that did not exist in baseline
Platform | Integration✅ Latency percentile monitoring (p95) alerts when exceeding 3 seconds for synchronous user-facing calls
Platform | Integration✅ Throughput monitoring alerts when above 80% of known rate limit or SLA threshold
Platform | Integration✅ Circuit breaker state monitoring alerts when any circuit opens indicating external system degradation
Platform | Integration✅ Retry rate monitoring alerts when sustained retry rate exceeds 5% indicating reliability issues
Platform | Integration✅ Integration telemetry routes to centralized monitoring platforms enabling correlation across system boundaries

Anti-Patterns

Where to lookWhat bad looks like
Platform | User Flows⚠️ Components are monitored independently showing healthy operation while 12% of users abandon critical flows with no visibility into why
Platform | User Flows⚠️ Component health dashboards show green while users cannot complete critical workflows revealing monitoring architecture failure
Platform | User Flows⚠️ User journey touching multiple Apex classes, flows, platform events, and external integrations lacks end-to-end visibility
Platform | User Flows⚠️ Critical business journeys lack instrumentation; success rates and abandonment points are unknown
Platform | Integration⚠️ Integration failures are detected only after cascading into user-facing failures rather than proactively
Platform | Integration⚠️ Integration telemetry remains isolated in individual systems preventing correlation and root cause analysis across boundaries
Platform | Integration⚠️ Monitoring alerts on downstream symptoms after damage is done rather than revealing root causes early

Patterns

Where to lookWhat good looks like
Platform | SLIs✅ Availability SLIs measure percentage of time critical flows execute successfully from user perspective (receiving expected result, not just 200 status code)
Platform | SLIs✅ Latency SLIs measure time to complete operations at p50 (median), p90 (90th percentile excluding outliers), p95 (95th percentile excluding worst outliers), and p99 (99th percentile showing worst-case experience)
Platform | SLIs✅ Latency targets vary by operation criticality and user expectation (500ms search feels instantaneous, same latency for typing input feels sluggish)
Platform | SLIs✅ Throughput SLIs measure volume of successful operations per time unit revealing capacity utilization and growth trends
Platform | SLIs✅ Throughput combined with latency reveals whether system maintains performance under increasing load or degrades as volume grows
Platform | SLIs✅ Error rate SLIs distinguish between error types (user errors like validation failures versus system errors like governor limit exceptions)
Platform | SLOs✅ Service Level Objectives are established defining acceptable SLI thresholds based on business requirements before selecting architecture patterns
Platform | SLOs✅ Business requirement “Sales reps must see opportunity updates within 5 seconds” drives SLO “Opportunity detail page loads complete within 3 seconds at p95” which drives architecture decisions
Platform | SLOs✅ SLOs emerge from business requirements and drive architectural decisions rather than being defined based on what system currently achieves

Anti-Patterns

Where to lookWhat bad looks like
Platform | SLIs⚠️ Availability is measured from technical perspective (200 status code) rather than user perspective (receiving expected result)
Platform | SLIs⚠️ Latency is measured only at median without understanding worst-case user experience at p95 or p99 percentiles
Platform | SLIs⚠️ Error rate SLIs do not distinguish between user errors and system errors, masking different problems requiring different solutions
Platform | SLOs⚠️ Architecture is built first, then SLO is defined as “whatever we currently achieve” leading to acceptance of 12-second page loads because system design never considered user requirements
Platform | SLOs⚠️ SLOs are architecture-driven rather than business-driven, setting targets based on what system achieves rather than what users need

Patterns

Where to lookWhat good looks like
Platform | Alerting✅ Severity classification routes critical incidents (production unavailable, data corruption risk, security compromise) to page on-call engineers immediately
Platform | Alerting✅ High severity incidents (degraded performance, partial outages) generate notifications for prompt investigation within business hours
Platform | Alerting✅ Warnings (approaching thresholds, configuration drift) generate daily digest notifications for proactive remediation
Platform | Alerting✅ Alert context includes affected user count, error messages, stack traces, recent deployments, related alerts, and links to relevant dashboards
Platform | Alerting✅ Rich context reduces mean time to acknowledge by enabling immediate severity assessment and investigation without gathering basic information
Platform | Alerting✅ Alert grouping prevents alert storms where hundreds of related alerts fire for single underlying problem
Platform | Alerting✅ Integration failure triggers one alert grouped by integration endpoint showing 156 failed requests in past 5 minutes with latency graph, error distribution, and dashboard link
Platform | Alerting✅ Alert suppression prevents known-transient conditions from generating alerts during expected maintenance windows or auto-remediation periods
Platform | Alerting✅ Alerts are suppressed during scheduled maintenance, during first few minutes after deployment while caches warm, and during brief transient failures where retry logic succeeds

Anti-Patterns

Where to lookWhat bad looks like
Platform | Alerting⚠️ All alerts route to same channel without severity classification; critical incidents are not distinguished from routine warnings
Platform | Alerting⚠️ Alerts lack diagnostic context; on-call engineers must gather basic information through separate queries before assessing severity
Platform | Alerting⚠️ Integration failure generates 156 individual alerts in 5 minutes overwhelming on-call engineer’s phone, each alert identical except for timestamp
Platform | Alerting⚠️ Alert storms prevent responders from distinguishing single underlying problem from hundreds of related symptoms
Platform | Alerting⚠️ Alerts fire during known maintenance windows and transient conditions generating noise that erodes trust in monitoring
Platform | Alerting⚠️ Alert fatigue from noisy false positives causes teams to ignore notifications reducing response effectiveness