Reliability - Disaster Recovery Patterns


Learn more about Well-Architected Reliability → Disaster Recovery and Business Continuity

Patterns

Where to lookWhat good looks like
Data backup strategy✅ Weekly or monthly automated exports using Data Export Service or a dedicated backup and recovery service (such as Own). Export frequency matches RPO requirements. Critical object data exported to external object storage where required.
Metadata backup strategy✅ All metadata versioned in Git using Salesforce DX source format. Metadata changes deployed only from source control. CI/CD pipeline can recreate any environment from repository. Point-in-time recovery available through Git history.
File backup strategy✅ ContentVersion records, attachments, and documents exported to external storage on schedule. Automated file export for regulatory retention requirements. External archival system for files older than business operational needs.
Backup validation✅ Quarterly restore drills to Full Sandbox or Developer environments. Restore procedures documented in runbooks with step-by-step instructions. Test restoration validates backup integrity and completeness before disasters occur.
Recovery testing schedule✅ Quarterly partial restore tests (metadata, specific objects, file subsets). Annual full recovery drill with complete environment restoration. Lessons learned documented and runbooks updated after every test.
Third-party backup service✅ Automated backup with point-in-time restore capability. Service-level agreement defining RPO/RTO. Daily automated verification of backup completion. Geographic redundancy with backups stored in multiple regions.
Custom objects and fields✅ Custom objects included in scheduled Data Export Service runs, with formula and roll-up summary field values recomputed rather than exported (no stored value exists to pull) and file or attachment binaries handled through a separate export path. Complex field types and lookup relationships documented and tested in the restore sequence.
Files and attachments✅ ContentDocument and ContentVersion objects exported with file binaries. Document libraries backed up including folder structures and permissions. Files exported maintain version history for critical document types.
Sandbox refresh strategy✅ Production org refreshed to Full Sandbox quarterly for load testing and disaster recovery validation. Sandbox refresh used as restore procedure test. Data masking applied post-refresh to protect sensitive data in test environments.
Backup monitoring and alerting✅ Automated alerts when scheduled backups fail or are delayed. Dashboard showing backup job success rate, completion time, and data volume trends. On-call escalation for backup failures exceeding acceptable delay window.

Anti-Patterns

Where to lookWhat bad looks like
Backup strategy⚠️ No automated backup configured, relying on platform infrastructure backups alone. Third-party backup service configured but never tested - backups may be corrupted or incomplete. Manual Data Loader exports relied on for backups (inconsistent execution, incomplete coverage, no automation).
Metadata versioning⚠️ Metadata changes deployed via Change Sets without source control. No Git repository for metadata history. Point-in-time recovery impossible - cannot restore last-known-good configuration after destructive changes.
Backup validation⚠️ Backups configured but never restored to test environment. First restore attempt occurs during actual disaster revealing incomplete scope or corrupted archives. Backup procedures documented but not tested by backup personnel.
Backup scope⚠️ Assuming Data Export Service captures everything when formula and roll-up summary field values never export (recalculated dynamically, not stored) and file or attachment binaries require separate handling, leaving gaps discovered only at restore.
Recovery procedures⚠️ Recovery runbook contains outdated steps from 2-year-old architecture. Recovery procedures tested by single expert who has since left company. Manual procedures so complex that execution under stress prone to errors.
Backup frequency⚠️ Critical transaction data backed up weekly when business requirement is 24-hour RPO. Backup frequency not aligned with business criticality - same schedule for all data regardless of importance. Backup schedule missed due to lack of monitoring and alerting.
Regulatory compliance⚠️ Backup retention shorter than the governing regulation requires. For example, HIPAA compliance documentation (policies, risk analyses, BAAs, breach records) purged before its six-year §164.316 retention, or SOX-governed financial audit records deleted before their seven-year retention. Compliance gap discovered during audit.
File storage⚠️ Large files stored in Salesforce without archival strategy. File storage grows until org approaches limit. No backup of ContentVersion records - files lost if org corrupted. Relying on users to maintain personal copies of critical files.
Sandbox testing⚠️ Full Sandbox never refreshed from production - testing environment wildly different from production. Data volume in sandbox 1% of production preventing realistic restore validation. Sandbox refresh treated as inconvenience rather than disaster recovery validation opportunity.
Monitoring⚠️ Backup jobs fail silently with no alerting. Discover backup failures weeks later during disaster when backup needed. No dashboard showing backup job success trends. Backup service health monitored only when invoice arrives.
Encryption keys⚠️ Treating Shield-encrypted data as safe without protecting key material: rotating keys alone leaves prior data intact, but destroying a tenant secret (crypto-shred) or losing Bring Your Own Key or cache-only key material renders the encrypted data permanently unrecoverable. Key management is omitted from recovery procedures.
Cross-org dependencies⚠️ Backing up the primary org alone when the solution depends on connected apps, external integrations, or SSO: connected app secrets, named-credential and OAuth authorizations, and external-auth tokens are not preserved by a data-and-metadata restore and must be reconfigured and re-authorized before a recovered org can authenticate to external systems. When external SSO is the sole login path, an unavailable identity provider blocks login even into a fully recovered org.

Patterns

Where to lookWhat good looks like
Recovery Time Objective (RTO) definition✅ RTO defined per business capability rather than uniform across entire org. Revenue-critical operations: minutes to 1 hour. Customer-facing services: 1-4 hours. Internal tools: 4-24 hours. Historical reporting: days. RTOs documented with business stakeholder approval.
Recovery Point Objective (RPO) definition✅ RPO defined based on acceptable data loss per capability. Financial transactions: near-zero RPO (seconds, event-driven async replication). Customer records: near-zero (minutes via CDC). Analytics data: hours (scheduled sync). Temporary workflow state: days (no replication needed).
Team availability and roles✅ Primary and backup personnel documented for all critical roles. No single point of failure in operational knowledge. Escalation tree defining primary responder, first backup, second backup. Contact information maintained in system accessible when Salesforce is unavailable.
Communication procedures✅ Incident communication plan defining who notifies users, customers, executives. External status page hosted outside Salesforce for user communication during platform outages. SMS notification system for on-call team when primary communication tools are unavailable. Escalation timeframes documented (when to escalate from team to manager to executive).
Vendor dependency mapping✅ All critical external vendors documented including Salesforce, integration partners, ISV package providers. Vendor SLAs and support windows documented (24/7 vs business-hours-only). Vendor escalation contacts and procedures documented. Alternative vendor options identified for critical dependencies.
Regulatory notification obligations✅ Mandatory incident notification requirements documented per industry (financial services, healthcare, government). Notification timeframes defined to match each governing regulation’s required window. Responsibility matrix defining who files notifications with which regulatory bodies. Templates prepared for rapid regulatory notification.
Business impact analysis✅ Each business capability ranked by revenue impact, user impact, and regulatory risk. Recovery prioritization based on business impact ranking. Resources allocated to disaster recovery proportional to business criticality. Impact analysis reviewed annually as business priorities evolve.

Anti-Patterns

Where to lookWhat bad looks like
RTO/RPO definition⚠️ “The system should be reliable” without measurable RTO/RPO targets. RTO/RPO defined by IT team without business stakeholder input. Uniform targets across all capabilities ignoring criticality differences (“everything needs 99.999% uptime”).
Team readiness⚠️ Single Salesforce admin knows disaster recovery procedures with no documentation or backup personnel. On-call rotation includes personnel who have never executed recovery procedures. Primary responders unavailable during disaster (vacation, illness, departed company).
Communication planning⚠️ No communication plan - discover need to notify users after incident already in progress. Status page hosted on Salesforce itself unavailable during Salesforce outage. Waiting hours to notify executives while attempting recovery.
Vendor dependencies⚠️ Unaware of vendor SLA terms until disaster occurs and expecting 24/7 support from vendor with business-hours-only coverage. No documented vendor escalation contacts. ISV package provider required for recovery but company no longer in business.
Regulatory requirements⚠️ Unaware of mandatory breach notification laws. Discover 24-hour notification requirement after 48 hours have passed. No templates prepared - drafting regulatory notifications during incident response while compliance clock ticking.
Recovery prioritization⚠️ Attempting to restore all capabilities simultaneously during disaster instead of prioritizing revenue-critical operations. Spending 8 hours recovering internal reporting capability before addressing customer-facing order system.
Testing and drills⚠️ Business continuity plan never tested through drills. Assuming documented procedures will work under stress without validation. Tabletop exercises avoided due to “lack of time” despite spending days recovering from undiscovered issues during actual disasters.
Knowledge transfer⚠️ Disaster recovery knowledge concentrated in one senior architect who documented nothing. New team members onboarded without disaster recovery training. Assuming everyone knows what to do during disasters without explicit training or drills.
Assumption validation⚠️ Recovery procedures assume backup accessible within 15 minutes but backup retrieval from cold storage takes 8 hours. Assuming team available within 30 minutes but no on-call rotation or escalation procedures defined. Recovery time estimates based on optimistic assumptions never validated through testing.

Patterns

Where to lookWhat good looks like
RTO architectural alignment✅ Architecture decisions match RTO commitments. Minutes RTO requires automated failover and hot standby. Hours RTO requires scripted recovery and warm standby. Days RTO accepts manual recovery from cold backup. Over-engineering avoided where business does not justify complexity.
RPO replication strategy✅ Replication method matches RPO target. Near-zero RPO: asynchronous event streaming via Platform Events on every record save (seconds of propagation lag, at-least-once delivery). Minutes RPO: Change Data Capture async replication. Hours RPO: scheduled batch API extract. Days RPO: manual backup on fixed schedule.
RTO testing✅ Recovery procedures tested under timed conditions. Actual RTO measured during disaster recovery drills. Gaps between target RTO and achieved RTO identified and remediated. RTO achievable by backup personnel, not just primary experts.
RPO validation✅ Data loss measured during recovery tests. Actual RPO compared to target during drills. Replication lag monitored continuously in production. Alerts configured when replication falls behind RPO commitment.
Capability-specific RTO/RPO✅ Different business capabilities have different RTO/RPO targets reflecting business criticality. E-commerce order flow: 15-minute RTO, near-zero RPO (justifies multi-org active-passive). Internal knowledge base: 8-hour RTO, 24-hour RPO (justifies manual restore from backup).
Cost-benefit analysis✅ RTO/RPO targets balanced against implementation cost. More stringent targets require exponentially increasing investment. Business case documented showing downtime cost justifying disaster recovery investment. Annual review confirms continued alignment between investment and business value.
SLA alignment✅ Solution RTO/RPO targets consider platform SLA baseline. Solution can only provide RTO better than platform baseline with multi-org architecture. Realistic targets set considering platform capabilities as foundation.
Monitoring of recovery capabilities✅ Continuous monitoring of backup currency (backup age vs RPO target). Alerting when backup falls behind schedule. Automated validation that recovery infrastructure (secondary org, external storage) is operational.
Automated vs manual recovery✅ Aggressive RTO targets (minutes) require automated failover with minimal human intervention. Moderate RTO targets (hours) allow scripted recovery with human initiation. Relaxed RTO targets (days) accept manual recovery following documented procedures.

Anti-Patterns

Where to lookWhat bad looks like
Unrealistic targets⚠️ Committing to 15-minute RTO with manual recovery procedures requiring 4 hours to execute. Promising zero RPO without synchronous replication infrastructure. RTO targets set by IT without architectural support.
Architecture misalignment⚠️ 30-minute RTO commitment but no automated failover and recovery procedures are entirely manual. Zero RPO commitment but only daily backups configured. Aggressive targets without corresponding investment in automation and redundancy.
Untested targets⚠️ RTO never validated through timed disaster recovery drills. Assuming 2-hour RTO achievable but never actually restoring under time constraints. Discovering during real disaster that documented 4-hour RTO actually requires 18 hours.
Uniform targets⚠️ All capabilities assigned same RTO/RPO regardless of business criticality. Internal batch reporting system designed for five-nines availability while revenue-critical e-commerce has same target. Massive over-investment in non-critical capabilities and under-investment in critical systems.
Ignoring dependencies⚠️ RTO assumes Salesforce org restoration only but solution depends on external identity provider also unavailable. RPO assumes data backup sufficient but integration credentials and configuration not backed up. Recovery time estimates ignore time to restore dependent systems.
Replication gaps⚠️ Committing to 15-minute RPO but Change Data Capture configured for only 3 of 12 critical objects. Platform Event replication configured but no monitoring of event delivery lag. Replication infrastructure breaks and goes undetected for weeks until disaster.
Cost ignorance⚠️ Committing to five-nines availability without understanding implementation cost. RTO/RPO targets set without cost-benefit analysis. Business unwilling to fund infrastructure required for targets already committed to customers.
Static targets⚠️ RTO/RPO defined 5 years ago never reassessed despite business growth and changing priorities. Targets continue unchanged as solution evolves and criticality changes. Annual review process absent.
Documentation gaps⚠️ RTO/RPO targets documented in architecture document but not communicated to business stakeholders. Business unaware of committed recovery capabilities. Gap discovered when business expects 1-hour recovery but architecture supports only 24-hour RTO.
False precision⚠️ Committing to “99.97% availability” when monitoring infrastructure cannot measure to that precision. Targets defined with false precision suggesting rigor when underlying analysis is rough estimation.
Monitoring absence⚠️ No continuous monitoring of recovery capabilities. Backup falls 2 days behind schedule violating RPO commitment but no alerts. Secondary org in active-passive architecture not receiving replication updates but health checks not configured.
Planning without execution⚠️ Beautiful disaster recovery plan with aggressive RTO/RPO targets but no automation, testing, or validated procedures. Plan created to satisfy compliance checkbox but no operational reality behind commitments.

Patterns

Where to lookWhat good looks like
Data redundancy✅ Platform provides infrastructure backup automatically. Application-level replication supplements platform backup when business requires faster recovery. Change Data Capture or Platform Events replicate critical data to secondary storage continuously. Replication enables recovery from logical corruption or configuration errors that infrastructure backups cannot address.
Application redundancy✅ Stateless application logic that survives individual transaction failures. No server-side state preventing horizontal scaling. Custom Metadata Types and Custom Settings for configuration instantly available across all application servers. Any application server can process any request without depending on specific server state.
Integration redundancy✅ Circuit breaker patterns detect failing integrations. Platform Events queue requests when external systems unavailable rather than blocking user operations. External system failures isolated from user-facing functionality. Multiple external vendor options identified for critical integrations enabling rapid provider switch during vendor outages.
Multi-org patterns for critical workloads✅ Active-passive pattern: primary org serves all traffic, secondary org synchronized but idle, failover during primary region outage. Active-active pattern: both orgs serve production traffic continuously, users allocated by geography or business unit. Multi-org justified only when business requires RPO/RTO beyond platform capabilities.
Multi-org data synchronization✅ Change Data Capture delivers near real-time replication for critical data (15-minute RPO). Platform Events provide event streaming for critical transactions (near-zero RPO). Bulk API 2.0 scheduled replication for reference data (24-hour RPO acceptable). Conflict resolution strategy defined for active-active when same record modified in both orgs.
Multi-org active-passive architecture✅ Primary org in US East region serves all production traffic. Secondary org in Europe region synchronized via CDC but idle. DNS routing or authentication layer directs users to active org. Automated health checks trigger failover when primary org unavailable for 10 consecutive minutes.
Multi-org active-active architecture✅ EMEA users on Europe org, AMER users on US org based on user profile region field. Bidirectional replication for shared reference data (accounts, products). Regional data stays in respective orgs (orders, cases). Conflict resolution uses last-write-wins with regional authority for customer master data.
Geographic distribution✅ Hyperforce regional data centers enable geographic distribution. Single-region deployment with platform-managed redundancy sufficient for 99.9% availability. Multi-org across regions considered only for 99.95%+ requirements or regulatory data residency mandates.
External system redundancy✅ Critical external APIs have primary and failover endpoints. Circuit breaker switches to failover endpoint when primary experiences sustained failures. External services with geographic redundancy preferred over single-region providers for critical integrations.
Infrastructure monitoring✅ trust.salesforce.com monitored via status API. Subscribe to instance-specific status notifications. Solution responds to platform health status dynamically by reducing non-critical load during degradation. Scale Center identifies resource-intensive operations before they impact redundancy capacity.
Failover automation✅ Automated failover preferred over manual for aggressive RTO targets. Health checks run every 1-5 minutes detecting failures. Failover decision logic avoids false positives (requires sustained failure, not transient errors). Automated rollback if failover destination also unhealthy.
Single points of failure✅ Architecture review identifies all single points of failure. Each identified SPOF assessed for business impact and mitigation cost. High-impact SPOFs mitigated through redundancy. Low-impact SPOFs accepted with documented risk acceptance and recovery procedures.
Testing redundancy capabilities✅ Scheduled failover drills validate automated failover logic. Backup systems tested under load (not just idle warm standby). Secondary org capacity validated handling production load during active-passive tests. Lessons learned from each test improve automation and procedures.

Anti-Patterns

Where to lookWhat bad looks like
Single points of failure⚠️ Entire solution depends on single external API with no fallback. Custom domain provider is SPOF - if provider fails, users cannot reach Experience Cloud. Single integration user account - if disabled, all integrations fail.
Untested redundancy⚠️ Multi-org active-passive architecture configured but never tested. Secondary org falls out of sync but health checks not configured. First failover attempt occurs during real disaster revealing secondary org corrupted or non-functional.
Data synchronization gaps⚠️ Active-passive multi-org with CDC replication configured for 5 of 30 critical objects. Failover occurs but secondary org missing 80% of required data. Replication lag monitoring absent - secondary org 3 days behind primary but no alerts.
False redundancy⚠️ Assuming multi-region architecture when both orgs actually hosted in same Hyperforce region. “Redundant” external APIs that actually use same upstream provider. Multiple integration paths that all route through same single point of failure.
Stateful architecture⚠️ Application logic depends on in-memory state preventing horizontal scaling. Single application server holds all user session state - server restart logs out all users. Static resources cached in application memory instead of Platform Cache creating inconsistency across servers.
No circuit breakers⚠️ External integration fails and every transaction waits full 120-second timeout. All users impacted when single external system unavailable. No failure isolation - one failing integration degrades entire platform performance.
Manual failover only⚠️ Multi-org architecture with only manual failover procedures. Requires 2-hour procedure to initiate failover while business is down. Automated health checks not implemented. Failover decision requires executive approval delaying recovery.
Capacity planning failure⚠️ Secondary org in active-passive architecture sized for 20% of production load. Failover occurs but secondary org cannot handle full production traffic. Active-active architecture with both orgs at 85% capacity leaving no redundancy when one org fails.
Configuration drift⚠️ Primary and secondary orgs diverge over time because deployments not synchronized. Metadata deployed to primary but not secondary. Failover reveals secondary org running 6-month-old code version.
Over-engineering⚠️ Multi-org architecture with active-active replication for internal admin tools used by 5 people. Operational complexity of multi-org not justified by business criticality. 99.999% availability target for batch reporting acceptable at 99%.
Geographic misunderstanding⚠️ Assuming Hyperforce “multi-region” provides multi-org capabilities automatically. Not understanding that single org is always in single region and multi-region requires multiple orgs. Geography-based redundancy impossible with single-org architecture.
Integration coupling⚠️ All integrations synchronous with no async fallback. External system unavailability blocks user operations. No queueing mechanism - transactions fail if external system temporarily down. Users cannot proceed when supporting integration unavailable.