Availability & Performance Patterns


Learn more about Well-Architected Reliability → High Availability Architecture

Patterns

Where to lookWhat good looks like
Business impact analysis✅ SLOs defined per business capability (e.g., “Opportunity close flow: 99.9% availability, p95 latency under 2 seconds”) rather than per technical component
User flow documentation✅ Each critical user flow has measurable SLO covering availability, latency, throughput, error rate, and recovery time tied to actual user experience
Architecture decision records✅ Availability targets based on business downtime impact analysis with documented rationale (e.g., internal batch reporting 99% vs revenue-critical order processing 99.95%)
SLO specifications✅ Realistic targets considering architectural complexity (99% = standard platform capabilities, 99.5% = basic redundancy and active monitoring, 99.9% = multi-region awareness and automated failover, 99.95% = active-active patterns and chaos testing, 99.99% = multi-org architecture)
Platform documentation✅ Solution SLOs less stringent than platform SLAs (e.g., platform 99.9% SLA → solution 99.5% SLO) providing error budget for application-layer failures
Monitoring configuration✅ Service level indicators (SLIs) objectively measurable via platform instrumentation: trust.salesforce.com uptime, Experience Cloud analytics, Event Monitoring API response time, AsyncApexJob completion
User experience metrics✅ SLOs reflect complete user journeys (e.g., “Case submission flow: all 5 steps complete within 8 seconds for p95 users”) not individual API uptime

Anti-Patterns

Where to lookWhat bad looks like
Architecture documentation⚠️ “The system should be fast and reliable” (not measurable, drives no architecture decisions)
SLO specifications⚠️ “Everything needs 99.999% uptime” (unjustified cost ignoring business criticality differences between revenue operations and internal admin tools)
Monitoring strategy⚠️ Infrastructure CPU monitoring as primary SLI (platform metric architects cannot act on, not reflecting user-facing experience)
Platform capacity planning⚠️ Solution SLO matches platform SLA exactly at 99.9% (no error budget for application-layer issues, integration failures, or planned maintenance)
SLO documentation⚠️ SLOs defined per technical component rather than per user-facing business capability (component availability necessary but insufficient for user experience reliability)

Patterns

Where to lookWhat good looks like
Regional deployment strategy✅ Single-region deployment with platform-managed Hyperforce redundancy (multiple availability zones, automatic infrastructure failover) provides 99.9% SLO for most solutions
Multi-org architecture✅ Multi-org patterns only when business requires guaranteed RPO/RTO beyond platform capabilities, regulatory geographic data isolation, or complete regional independence (active-passive or active-active with data synchronization)
Data replication✅ Change Data Capture for critical data (near real-time RPO, minutes), Platform Events for near real-time business events, scheduled Bulk API 2.0 extraction for reference data (24-hour RPO)
Platform health monitoring✅ Subscription to trust.salesforce.com status notifications and instance-specific status APIs enabling proactive response to incidents, maintenance windows, and performance impacts
Dynamic load management✅ Solutions respond to platform health status by reducing non-critical batch processing during degradation, deferring background jobs during maintenance windows
Scale Center usage✅ Weekly Scale Center review identifies long-running transactions and operations consuming disproportionate resources, enabling architectural remediation before user impact
Failure detection layers✅ Multi-level detection: platform failures via status API, integration failures via timeout monitoring, application failures via exception logging, performance degradation via latency percentiles, capacity warnings via Proactive Monitoring
Alert thresholds✅ Alerts trigger on error rate patterns (e.g., 5 API timeouts in 60 seconds) or sustained degradation, not isolated failures that are normal in distributed systems

Anti-Patterns

Where to lookWhat bad looks like
Multi-org design⚠️ Multi-org architecture implemented “just in case” without documented business requirement justifying operational complexity of data synchronization, user provisioning, deployment coordination
Platform health monitoring⚠️ Ignore trust.salesforce.com status notifications, discover maintenance windows when users report failures (reactive rather than proactive response)
Failure detection⚠️ Alert on every individual error regardless of rate (generates alert fatigue obscuring critical signals); single API timeout triggers page-out (over-reaction to normal distributed system behavior)
Data replication⚠️ No integration-specific monitoring, only discover integration failures when users report blocked operations (missing visibility into external system health)
Redundancy strategy⚠️ Synchronous callout to external enrichment service with no timeout or fallback (user blocked when external service slow or unavailable)

Patterns

Where to lookWhat good looks like
Real User Monitoring (RUM)✅ Actual user latency measured via Experience Cloud analytics or custom instrumentation reflecting real network conditions, device performance, geographic distribution
Synthetic monitoring✅ Automated transactions from multiple locations validating availability and performance before users report issues (e.g., order submission every 15 minutes from 3 regions)
Transaction tracing✅ Multi-step operations instrumented with timing per step identifying bottlenecks (e.g., step 3 address validation callout accounts for 70% of total latency)
Query optimization✅ Selective queries with indexed filters, relationship queries reducing round trips, query optimizer consideration with production-scale data volumes (behavior changes dramatically with 10M+ records)
Platform Cache utilization✅ Session or org cache partitions serving frequently accessed reference data, pre-populated during successful operations providing stale-but-available data when real-time API unavailable
Timeout configuration✅ User-facing synchronous callouts 5-10 seconds (responsive UI), background async callouts 30-60 seconds (variable external performance), batch processing 120 seconds (no user waiting); budget total across multiple callouts per transaction
Integration latency monitoring✅ p50/p95/p99 latency percentiles per integration with SLO per endpoint; latency anomalies (response times drifting upward) indicate capacity saturation before complete failure

Anti-Patterns

Where to lookWhat bad looks like
Monitoring strategy⚠️ Monitor only infrastructure metrics (server CPU, network bandwidth) without user-facing flow measurement (metrics don’t reflect actual user experience)
Cache strategy⚠️ No Platform Cache utilization for frequently accessed reference data (repeated queries for same data increasing latency and consuming SOQL limits)
Timeout configuration⚠️ Synchronous Flow callout with 120-second timeout causing user to stare at loading spinner for 2 minutes (should use 5-10 second timeout with fallback)
Performance testing⚠️ Small-scale testing with minimal data volumes (query optimizer behavior changes dramatically at 10M+ records making small-scale results misleading)
Integration design⚠️ No transaction-level tracing of multi-step workflows (bottleneck step invisible in aggregate metrics)

Patterns

Where to lookWhat good looks like
Horizontal scaling patterns✅ Workload distributed across multiple resources: bulkification processes multiple records per transaction, asynchronous processing distributes work across time, data partitioning distributes load across boundaries
Batch processing configuration✅ Optimal batch size balancing throughput and governor limits (2,000 records per Batch Apex execute, 200 records per trigger invocation, 25 requests per Composite API call)
API rate management✅ Integration batching and caching reduces API consumption, monitors 24-hour API allocation by edition, capacity planning adds licenses or limit increases before exhaustion
Concurrent execution✅ Parallel batch jobs for partitioned data (5 regional rollup batches running simultaneously), Platform Event subscribers processing streams concurrently
Queue depth monitoring✅ For async integrations, queue depth trends indicate processing falling behind production rate requiring capacity increase or throughput optimization
Volume anomaly detection✅ Transaction volumes significantly above/below expected patterns may indicate runaway processes requiring investigation and throttling

Anti-Patterns

Where to lookWhat bad looks like
Batch processing⚠️ Process all records sequentially when data volume requires partitioned parallel execution (date-based, record type-based, or owner-based partitioning enables concurrent processing)
API consumption⚠️ Integration pattern adds 25% API consumption without capacity planning (sudden “API limit exceeded” errors impact users)
Queue monitoring⚠️ No monitoring of async integration queue depth (processing falls behind production rate without visibility)
Governor limit management⚠️ Synchronous operations processing large volumes hitting heap size or CPU time limits (should move to asynchronous processing with dedicated limits)

Patterns

Where to lookWhat good looks like
Feature criticality hierarchy✅ Defined tiers: Critical (revenue/compliance, never degraded), Important (core workflows, degraded only during major incidents), Supporting (enhanced experience, disabled during integration failure), Optional (nice-to-have, disabled during high load)
Circuit breaker implementation✅ Circuit states (Closed = normal, Open = fails immediately after threshold, Half-open = recovery probe); Platform Cache stores state, Platform Events broadcast changes; opens after consecutive failures (e.g., 10 failures)
Retry logic✅ Exponential backoff for transient failures (1s, 2s, 4s, 8s, cap at 30-60s); retry network timeouts and 5xx errors with backoff, retry 429 rate limits after Retry-After header; never retry client 4xx errors or governor limit errors in same transaction
Fallback strategies✅ Alternative data sources (Platform Cache when real-time API unavailable), default behavior (standard rules when personalization service down), manual process (admin interface for stuck transactions), queue for retry (Platform Events with 72h replay)
Error handling architecture✅ Fail fast (validate inputs at entry), fail gracefully (process 197 of 200 records when 3 fail validation), fail informatively (log with transaction ID, user context, stack trace), fail safely (rollback partial transactions, never expose internals to users)
Recovery automation✅ Platform monitoring triggers automated recovery: asynchronous retry mechanisms, dead letter queues for failed processing, self-healing workflows enabling rapid recovery without manual intervention

Anti-Patterns

Where to lookWhat bad looks like
Integration dependency⚠️ Order submission fails completely when external address validation service down (critical flow blocked by supporting feature; should degrade gracefully with manual review)
Circuit breaker strategy⚠️ No circuit breaker on external callout causing every transaction to wait 120 seconds for timeout when external system down (all users impacted by one failing integration)
Retry logic⚠️ Retry Apex CPU limit exceeded error immediately in same transaction context (will fail again; should re-queue as async operation with dedicated limits)
Retry logic⚠️ Retry indefinitely until success (infinite loop risk, transaction never completes; should cap maximum retry attempts)
Error handling⚠️ Try-catch block swallows exception with no logging (failure invisible to operators preventing diagnosis and remediation)
Error handling⚠️ Display stack trace to end user in Lightning page (security risk revealing implementation details, poor user experience)
Fallback strategy⚠️ Tax calculation service fails blocking all order submissions (critical capability with no fallback; should queue for manual processing or use default rates)

Patterns

Where to lookWhat good looks like
Health modeling✅ Aggregate health status: Service health (transaction success rate >99.5% green, 98-99.5% yellow, <98% red), Integration health (all responding green, degraded yellow, circuit breaker open red), Data health (sync current green, behind yellow, stale/failed red), Capacity health (<70% green, 70-85% yellow, >85% red)
Event Monitoring✅ EventLogFile objects with 24h/1h delivery routed to external SIEM for retention beyond native retention limits; captures API calls, page views, report exports, login activity, Apex execution
Proactive Monitoring alerts✅ Continuous evaluation surfaces API request limit spikes (75% threshold triggers investigation), concurrent Apex execution failures, SOQL row limit issues, storage consumption trends
Integration health dashboard✅ Per-integration metrics: error rate (<1% target), latency p50/p95/p99 with SLO, timeout rate (<0.1% target), circuit breaker state, async queue depth; structured logs with request ID, endpoint, response code, duration
Actionable alert design✅ Every alert defines who responds, what they check, how they remediate; includes threshold crossed, current value, recent trend, dashboard link, runbook link; severity matched to urgency (CRITICAL pages on-call for user impact, WARNING emails for trending concerns)
Anomaly detection✅ Volume anomalies (300% of typical load), error rate anomalies (2% vs 0.5% typical even if below 5% threshold), latency drift (gradual upward trend), behavioral anomalies (unusual login patterns, batch jobs outside schedule); used to guide investigation not trigger immediate escalation

Anti-Patterns

Where to lookWhat bad looks like
Event Monitoring⚠️ Event Monitoring enabled but never exported (native retention insufficient for trend analysis, historical data lost)
Alert design⚠️ Alert message “Error occurred” with no context, threshold, or response action (generates alert fatigue, operators can’t triage effectively)
Alert severity⚠️ Treat all anomalies as critical alerts (high false positive rate creates alert fatigue causing critical alerts to be missed)
Health monitoring⚠️ No aggregate health model, operators overwhelmed with individual technical metrics missing cross-system patterns
Integration monitoring⚠️ No integration-specific error rate, latency, or circuit breaker state tracking (integration failures discovered reactively through user reports)
Capacity monitoring⚠️ No Proactive Monitoring or Scale Center review (reliability risks from resource-intensive transactions go undetected until production impact)

For inline examples and detailed guidance, see the Reliability pillar, or browse the Patterns and Anti-Patterns explorer.