Availability & Performance Patterns
Learn more about Well-Architected Reliability → High Availability Architecture
Patterns
| Where to look | What good looks like |
|---|---|
| Business impact analysis | ✅ SLOs defined per business capability (e.g., “Opportunity close flow: 99.9% availability, p95 latency under 2 seconds”) rather than per technical component |
| User flow documentation | ✅ Each critical user flow has measurable SLO covering availability, latency, throughput, error rate, and recovery time tied to actual user experience |
| Architecture decision records | ✅ Availability targets based on business downtime impact analysis with documented rationale (e.g., internal batch reporting 99% vs revenue-critical order processing 99.95%) |
| SLO specifications | ✅ Realistic targets considering architectural complexity (99% = standard platform capabilities, 99.5% = basic redundancy and active monitoring, 99.9% = multi-region awareness and automated failover, 99.95% = active-active patterns and chaos testing, 99.99% = multi-org architecture) |
| Platform documentation | ✅ Solution SLOs less stringent than platform SLAs (e.g., platform 99.9% SLA → solution 99.5% SLO) providing error budget for application-layer failures |
| Monitoring configuration | ✅ Service level indicators (SLIs) objectively measurable via platform instrumentation: trust.salesforce.com uptime, Experience Cloud analytics, Event Monitoring API response time, AsyncApexJob completion |
| User experience metrics | ✅ SLOs reflect complete user journeys (e.g., “Case submission flow: all 5 steps complete within 8 seconds for p95 users”) not individual API uptime |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Architecture documentation | ⚠️ “The system should be fast and reliable” (not measurable, drives no architecture decisions) |
| SLO specifications | ⚠️ “Everything needs 99.999% uptime” (unjustified cost ignoring business criticality differences between revenue operations and internal admin tools) |
| Monitoring strategy | ⚠️ Infrastructure CPU monitoring as primary SLI (platform metric architects cannot act on, not reflecting user-facing experience) |
| Platform capacity planning | ⚠️ Solution SLO matches platform SLA exactly at 99.9% (no error budget for application-layer issues, integration failures, or planned maintenance) |
| SLO documentation | ⚠️ SLOs defined per technical component rather than per user-facing business capability (component availability necessary but insufficient for user experience reliability) |
Patterns
| Where to look | What good looks like |
|---|---|
| Regional deployment strategy | ✅ Single-region deployment with platform-managed Hyperforce redundancy (multiple availability zones, automatic infrastructure failover) provides 99.9% SLO for most solutions |
| Multi-org architecture | ✅ Multi-org patterns only when business requires guaranteed RPO/RTO beyond platform capabilities, regulatory geographic data isolation, or complete regional independence (active-passive or active-active with data synchronization) |
| Data replication | ✅ Change Data Capture for critical data (near real-time RPO, minutes), Platform Events for near real-time business events, scheduled Bulk API 2.0 extraction for reference data (24-hour RPO) |
| Platform health monitoring | ✅ Subscription to trust.salesforce.com status notifications and instance-specific status APIs enabling proactive response to incidents, maintenance windows, and performance impacts |
| Dynamic load management | ✅ Solutions respond to platform health status by reducing non-critical batch processing during degradation, deferring background jobs during maintenance windows |
| Scale Center usage | ✅ Weekly Scale Center review identifies long-running transactions and operations consuming disproportionate resources, enabling architectural remediation before user impact |
| Failure detection layers | ✅ Multi-level detection: platform failures via status API, integration failures via timeout monitoring, application failures via exception logging, performance degradation via latency percentiles, capacity warnings via Proactive Monitoring |
| Alert thresholds | ✅ Alerts trigger on error rate patterns (e.g., 5 API timeouts in 60 seconds) or sustained degradation, not isolated failures that are normal in distributed systems |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Multi-org design | ⚠️ Multi-org architecture implemented “just in case” without documented business requirement justifying operational complexity of data synchronization, user provisioning, deployment coordination |
| Platform health monitoring | ⚠️ Ignore trust.salesforce.com status notifications, discover maintenance windows when users report failures (reactive rather than proactive response) |
| Failure detection | ⚠️ Alert on every individual error regardless of rate (generates alert fatigue obscuring critical signals); single API timeout triggers page-out (over-reaction to normal distributed system behavior) |
| Data replication | ⚠️ No integration-specific monitoring, only discover integration failures when users report blocked operations (missing visibility into external system health) |
| Redundancy strategy | ⚠️ Synchronous callout to external enrichment service with no timeout or fallback (user blocked when external service slow or unavailable) |
Patterns
| Where to look | What good looks like |
|---|---|
| Real User Monitoring (RUM) | ✅ Actual user latency measured via Experience Cloud analytics or custom instrumentation reflecting real network conditions, device performance, geographic distribution |
| Synthetic monitoring | ✅ Automated transactions from multiple locations validating availability and performance before users report issues (e.g., order submission every 15 minutes from 3 regions) |
| Transaction tracing | ✅ Multi-step operations instrumented with timing per step identifying bottlenecks (e.g., step 3 address validation callout accounts for 70% of total latency) |
| Query optimization | ✅ Selective queries with indexed filters, relationship queries reducing round trips, query optimizer consideration with production-scale data volumes (behavior changes dramatically with 10M+ records) |
| Platform Cache utilization | ✅ Session or org cache partitions serving frequently accessed reference data, pre-populated during successful operations providing stale-but-available data when real-time API unavailable |
| Timeout configuration | ✅ User-facing synchronous callouts 5-10 seconds (responsive UI), background async callouts 30-60 seconds (variable external performance), batch processing 120 seconds (no user waiting); budget total across multiple callouts per transaction |
| Integration latency monitoring | ✅ p50/p95/p99 latency percentiles per integration with SLO per endpoint; latency anomalies (response times drifting upward) indicate capacity saturation before complete failure |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Monitoring strategy | ⚠️ Monitor only infrastructure metrics (server CPU, network bandwidth) without user-facing flow measurement (metrics don’t reflect actual user experience) |
| Cache strategy | ⚠️ No Platform Cache utilization for frequently accessed reference data (repeated queries for same data increasing latency and consuming SOQL limits) |
| Timeout configuration | ⚠️ Synchronous Flow callout with 120-second timeout causing user to stare at loading spinner for 2 minutes (should use 5-10 second timeout with fallback) |
| Performance testing | ⚠️ Small-scale testing with minimal data volumes (query optimizer behavior changes dramatically at 10M+ records making small-scale results misleading) |
| Integration design | ⚠️ No transaction-level tracing of multi-step workflows (bottleneck step invisible in aggregate metrics) |
Patterns
| Where to look | What good looks like |
|---|---|
| Horizontal scaling patterns | ✅ Workload distributed across multiple resources: bulkification processes multiple records per transaction, asynchronous processing distributes work across time, data partitioning distributes load across boundaries |
| Batch processing configuration | ✅ Optimal batch size balancing throughput and governor limits (2,000 records per Batch Apex execute, 200 records per trigger invocation, 25 requests per Composite API call) |
| API rate management | ✅ Integration batching and caching reduces API consumption, monitors 24-hour API allocation by edition, capacity planning adds licenses or limit increases before exhaustion |
| Concurrent execution | ✅ Parallel batch jobs for partitioned data (5 regional rollup batches running simultaneously), Platform Event subscribers processing streams concurrently |
| Queue depth monitoring | ✅ For async integrations, queue depth trends indicate processing falling behind production rate requiring capacity increase or throughput optimization |
| Volume anomaly detection | ✅ Transaction volumes significantly above/below expected patterns may indicate runaway processes requiring investigation and throttling |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Batch processing | ⚠️ Process all records sequentially when data volume requires partitioned parallel execution (date-based, record type-based, or owner-based partitioning enables concurrent processing) |
| API consumption | ⚠️ Integration pattern adds 25% API consumption without capacity planning (sudden “API limit exceeded” errors impact users) |
| Queue monitoring | ⚠️ No monitoring of async integration queue depth (processing falls behind production rate without visibility) |
| Governor limit management | ⚠️ Synchronous operations processing large volumes hitting heap size or CPU time limits (should move to asynchronous processing with dedicated limits) |
Patterns
| Where to look | What good looks like |
|---|---|
| Feature criticality hierarchy | ✅ Defined tiers: Critical (revenue/compliance, never degraded), Important (core workflows, degraded only during major incidents), Supporting (enhanced experience, disabled during integration failure), Optional (nice-to-have, disabled during high load) |
| Circuit breaker implementation | ✅ Circuit states (Closed = normal, Open = fails immediately after threshold, Half-open = recovery probe); Platform Cache stores state, Platform Events broadcast changes; opens after consecutive failures (e.g., 10 failures) |
| Retry logic | ✅ Exponential backoff for transient failures (1s, 2s, 4s, 8s, cap at 30-60s); retry network timeouts and 5xx errors with backoff, retry 429 rate limits after Retry-After header; never retry client 4xx errors or governor limit errors in same transaction |
| Fallback strategies | ✅ Alternative data sources (Platform Cache when real-time API unavailable), default behavior (standard rules when personalization service down), manual process (admin interface for stuck transactions), queue for retry (Platform Events with 72h replay) |
| Error handling architecture | ✅ Fail fast (validate inputs at entry), fail gracefully (process 197 of 200 records when 3 fail validation), fail informatively (log with transaction ID, user context, stack trace), fail safely (rollback partial transactions, never expose internals to users) |
| Recovery automation | ✅ Platform monitoring triggers automated recovery: asynchronous retry mechanisms, dead letter queues for failed processing, self-healing workflows enabling rapid recovery without manual intervention |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Integration dependency | ⚠️ Order submission fails completely when external address validation service down (critical flow blocked by supporting feature; should degrade gracefully with manual review) |
| Circuit breaker strategy | ⚠️ No circuit breaker on external callout causing every transaction to wait 120 seconds for timeout when external system down (all users impacted by one failing integration) |
| Retry logic | ⚠️ Retry Apex CPU limit exceeded error immediately in same transaction context (will fail again; should re-queue as async operation with dedicated limits) |
| Retry logic | ⚠️ Retry indefinitely until success (infinite loop risk, transaction never completes; should cap maximum retry attempts) |
| Error handling | ⚠️ Try-catch block swallows exception with no logging (failure invisible to operators preventing diagnosis and remediation) |
| Error handling | ⚠️ Display stack trace to end user in Lightning page (security risk revealing implementation details, poor user experience) |
| Fallback strategy | ⚠️ Tax calculation service fails blocking all order submissions (critical capability with no fallback; should queue for manual processing or use default rates) |
Patterns
| Where to look | What good looks like |
|---|---|
| Health modeling | ✅ Aggregate health status: Service health (transaction success rate >99.5% green, 98-99.5% yellow, <98% red), Integration health (all responding green, degraded yellow, circuit breaker open red), Data health (sync current green, behind yellow, stale/failed red), Capacity health (<70% green, 70-85% yellow, >85% red) |
| Event Monitoring | ✅ EventLogFile objects with 24h/1h delivery routed to external SIEM for retention beyond native retention limits; captures API calls, page views, report exports, login activity, Apex execution |
| Proactive Monitoring alerts | ✅ Continuous evaluation surfaces API request limit spikes (75% threshold triggers investigation), concurrent Apex execution failures, SOQL row limit issues, storage consumption trends |
| Integration health dashboard | ✅ Per-integration metrics: error rate (<1% target), latency p50/p95/p99 with SLO, timeout rate (<0.1% target), circuit breaker state, async queue depth; structured logs with request ID, endpoint, response code, duration |
| Actionable alert design | ✅ Every alert defines who responds, what they check, how they remediate; includes threshold crossed, current value, recent trend, dashboard link, runbook link; severity matched to urgency (CRITICAL pages on-call for user impact, WARNING emails for trending concerns) |
| Anomaly detection | ✅ Volume anomalies (300% of typical load), error rate anomalies (2% vs 0.5% typical even if below 5% threshold), latency drift (gradual upward trend), behavioral anomalies (unusual login patterns, batch jobs outside schedule); used to guide investigation not trigger immediate escalation |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Event Monitoring | ⚠️ Event Monitoring enabled but never exported (native retention insufficient for trend analysis, historical data lost) |
| Alert design | ⚠️ Alert message “Error occurred” with no context, threshold, or response action (generates alert fatigue, operators can’t triage effectively) |
| Alert severity | ⚠️ Treat all anomalies as critical alerts (high false positive rate creates alert fatigue causing critical alerts to be missed) |
| Health monitoring | ⚠️ No aggregate health model, operators overwhelmed with individual technical metrics missing cross-system patterns |
| Integration monitoring | ⚠️ No integration-specific error rate, latency, or circuit breaker state tracking (integration failures discovered reactively through user reports) |
| Capacity monitoring | ⚠️ No Proactive Monitoring or Scale Center review (reliability risks from resource-intensive transactions go undetected until production impact) |
For inline examples and detailed guidance, see the Reliability pillar, or browse the Patterns and Anti-Patterns explorer.