Incident Response Patterns
Learn more about Well-Architected Trust → Security Incident Response
Security incidents will occur despite preventive measures. Trust architecture requires designing for detectability, rapid response, forensic investigation, and graceful recovery. This page consolidates incident response patterns from the Trust pillar’s Incident Response section, covering detection architecture, response procedures, forensic capabilities, and post-incident learning.
Patterns
| Where to look | What good looks like |
|---|---|
| Event Monitoring | ✅ Enable Event Monitoring capturing login patterns, bulk data exports, permission changes, and API activity with routing to SIEM platforms for retention exceeding native retention limits |
| Transaction Security Policies | ✅ Configure policies evaluating events in real time to block suspicious actions, require step-up authentication, or trigger notifications based on behavioral baselines |
| Setup Audit Trail | ✅ Monitor administrative configuration changes enabling detection of unauthorized modifications with 180-day retention |
| Custom Application Logging | ✅ Implement security-relevant event logging in Apex capturing authentication failures, authorization denials, and suspicious input patterns |
| Health Check Monitoring | ✅ Configure automated detection of configuration drift from security baselines that might indicate unauthorized changes |
| SIEM Integration | ✅ Route Event Monitoring logs to enterprise SIEM platforms for correlation with enterprise security telemetry and behavioral analysis |
| Alert Rules | ✅ Design alert rules detecting suspicious patterns (bulk access outside business hours, login from new geographies, rapid permission escalation) with behavioral baselines to minimize false positives |
| Behavioral Analytics | ✅ Establish behavioral baselines enabling anomaly detection for unusual data access volumes, off-hours administrative activity, and API consumption patterns |
| Multi-Channel Detection | ✅ Implement detection through Event Monitoring, Transaction Security, Setup Audit Trail, custom logging, and Health Check providing multiple independent detection signals |
| Immutable Log Storage | ✅ Route logs to immutable external storage that attackers cannot modify even with administrative Salesforce access, ensuring forensic evidence integrity |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Native Log Retention Only | ⚠️ Relying solely on native Salesforce audit trail retention windows (Event Monitoring native limits, 180 days Setup Audit Trail) without external archiving, discovering logs expired when investigation requires historical analysis |
| No Behavioral Baselines | ⚠️ Having no behavioral baselines for anomaly detection, generating excessive false positive alerts that train security teams to ignore notifications |
| No SIEM Integration | ⚠️ Lacking enterprise SIEM integration, missing correlation opportunities with security telemetry from other systems that would reveal broader compromise patterns |
| Manual Log Review | ⚠️ Depending on periodic manual log review rather than automated alerting, discovering incidents weeks or months after occurrence when damage is extensive |
| No Transaction Security Policies | ⚠️ Having no real-time Transaction Security policies, lacking capability to block suspicious actions preventing damage before it occurs |
| Incomplete Coverage | ⚠️ Monitoring only subset of security-relevant events (login but not data access, configuration changes but not permission modifications) creating blind spots |
| Alert Fatigue | ⚠️ Generating excessive false positive alerts without tuning, training teams to ignore notifications including legitimate security events |
Patterns
| Where to look | What good looks like |
|---|---|
| Isolation Boundaries | ✅ Design architecture so compromised components can be isolated through permission set revocation, IP restriction changes, or session termination without disrupting critical business functions |
| Session Management | ✅ Implement rapid session termination capabilities for compromised accounts including force logout and session invalidation across all active sessions |
| Permission Revocation | ✅ Grant privileges through permission sets rather than baseline profiles so revocation is rapid: removing the permission set assignment revokes access when no profile or other assignment independently grants the same permission |
| IP Restrictions | ✅ Configure profile-level Login IP Ranges for privileged accounts so logins from other addresses are denied outright, enabling rapid network-level isolation to known administrative locations |
| User Account Freezing | ✅ Implement account freezing procedures that prevent authentication while preserving data for investigation without deleting the compromised account |
| API Access Controls | ✅ Design API authentication with token revocation capabilities enabling rapid isolation of compromised integration accounts |
| Multi-Org Isolation | ✅ For high-value assets, evaluate multi-org architecture providing strongest isolation for critical components at cost of increased operational complexity |
| Emergency Break-Glass Procedures | ✅ Maintain documented break-glass access procedures for emergency administrative access with post-use audit and validation |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| No Isolation Boundaries | ⚠️ Lacking architectural isolation boundaries requiring complete system shutdown to contain any compromise, disrupting all business functions |
| Overly Permissive Access | ⚠️ Granting broad Modify All Data permissions preventing granular privilege revocation without disrupting multiple business processes |
| Shared Service Accounts | ⚠️ Using single shared “API User” or “Integration User” account across multiple systems preventing isolation of specific compromised integration without breaking all integrations |
| No IP Restrictions | ⚠️ Having no IP allowlisting for privileged accounts, lacking network-level containment capabilities when accounts are compromised |
| No Session Management | ⚠️ Lacking session termination capabilities requiring password reset as only way to force logout, introducing delays during critical response windows |
| No Emergency Procedures | ⚠️ Having no documented break-glass or emergency access procedures, improvising critical decisions under pressure during active incidents |
Patterns
| Where to look | What good looks like |
|---|---|
| Event Monitoring Logs | ✅ Capture detailed activity logs including user actions, API calls, authentication events, and data access patterns with sufficient detail for forensic analysis |
| Field Audit Trail | ✅ Enable Field Audit Trail preserving data change history for Restricted and Confidential fields with retention matching regulatory requirements (supports indefinite retention; confirm the required period against the governing regulation) |
| Immutable Storage | ✅ Route logs to external immutable storage that attackers cannot modify, ensuring forensic evidence integrity throughout investigation |
| Setup Audit Trail | ✅ Preserve administrative configuration change history enabling investigation of unauthorized system modifications with 180-day retention |
| Scope Assessment | ✅ Design permission models and data access controls enabling rapid assessment of data exposure scope based on compromised account privileges |
| Change Correlation | ✅ Maintain correlation between Setup Audit Trail configuration changes and Event Monitoring activity patterns enabling investigation of attacker techniques |
| Data Access Logging | ✅ Log all access to Restricted and Confidential data enabling investigation of what data was exposed during compromise period |
| Integration Audit Trail | ✅ Capture API authentication, authorization, and data exchange patterns for integrated systems enabling investigation of cross-system compromise |
| Evidence Preservation | ✅ Route logs and audit trails continuously to immutable external storage so forensic evidence is preserved and available when an incident is investigated |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Mutable Log Storage | ⚠️ Storing logs only within Salesforce where attackers with administrative access can delete or modify forensic evidence |
| Insufficient Log Detail | ⚠️ Capturing high-level events without sufficient detail for investigation (login recorded but not IP address, data access recorded but not specific records) |
| No Data Access Logging | ⚠️ Lacking detailed logging of access to Restricted and Confidential data, unable to assess data exposure scope during breach notification |
| Missing Field Audit Trail | ⚠️ Not enabling Field Audit Trail for sensitive fields, unable to determine what data values were exposed during compromise |
| No Change Correlation | ⚠️ Having no correlation capability between Setup Audit Trail configuration changes and subsequent Event Monitoring activity patterns |
| Deleted Audit History | ⚠️ Allowing audit history to expire before the regulatory retention minimum the governing regulation sets for your deployment, creating compliance violations |
Patterns
| Where to look | What good looks like |
|---|---|
| Recovery Time Objectives | ✅ Define and test Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for security incident scenarios, not just disaster recovery scenarios |
| Tested Recovery Procedures | ✅ Document recovery processes for compromised accounts, bulk data exfiltration, unauthorized deployments, privilege escalation, and package compromise with quarterly testing |
| Backup Validation | ✅ Validate backup integrity regularly ensuring restoration capabilities function during security incidents, not just operational failures |
| Rollback Capabilities | ✅ Design deployment pipelines with tested rollback procedures for unauthorized code deployments or malicious configuration changes |
| Credential Rotation | ✅ Implement rapid credential rotation procedures for compromised integration accounts, API users, and named credentials |
| Permission Restoration | ✅ Document procedures for restoring appropriate access after overly broad isolation actions during incident response |
| Data Restoration | ✅ Design data recovery capabilities addressing data corruption or deletion resulting from security incidents, distinct from disaster recovery |
| Service Continuity | ✅ Architect solutions so security incident response procedures maintain critical business function continuity during investigation and remediation |
| Restoration Validation | ✅ Test recovery procedures quarterly validating that restored systems function correctly and compromise has not persisted through restoration |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Untested Procedures | ⚠️ Having documented recovery procedures never tested through exercises, discovering during actual incidents that procedures are incomplete or incorrect |
| No RTO/RPO for Security | ⚠️ Defining Recovery Time Objectives only for disaster recovery without considering security incident scenarios requiring different recovery procedures |
| No Rollback Capability | ⚠️ Lacking deployment rollback procedures for unauthorized code deployments, requiring full redevelopment to remove malicious changes |
| Unvalidated Backups | ⚠️ Never testing backup restoration, discovering during incidents that backups are corrupted or incomplete |
| No Credential Rotation Procedures | ⚠️ Having no rapid credential rotation procedures for compromised integration accounts, leaving compromised credentials active during investigation |
| Manual Recovery Processes | ⚠️ Depending entirely on manual recovery procedures during incidents, introducing errors and delays when speed is critical |
| No Restoration Testing | ⚠️ Never validating that restored systems function correctly after recovery, potentially reintroducing compromise or data corruption |
Patterns
| Where to look | What good looks like |
|---|---|
| Compromised User Account Response | ✅ Document procedures: freeze user, force credential reset, review Setup Audit Trail and data access logs for compromise period, assess data exposure scope, execute breach notification if Restricted data accessed |
| Bulk Data Exfiltration Response | ✅ Document procedures: revoke session immediately, restrict permissions, identify affected records and classification levels, execute regulatory breach notification within the timeline the governing regulation sets for your deployment |
| Unauthorized Code Deployment Response | ✅ Document procedures: rollback deployment immediately, audit all changes from compromised credential, review pipeline security for authorization bypass, scan deployed code for persistent backdoors |
| Privilege Escalation Response | ✅ Document procedures: revoke escalated privileges immediately, audit activities performed with elevated access, investigate how escalation bypassed approval controls, implement compensating controls |
| AppExchange Package Compromise Response | ✅ Document procedures: isolate package permissions, assess data accessed by package, evaluate package replacement options, document lessons learned for package evaluation criteria |
| Detection Signals | ✅ Configure monitoring detecting each scenario: login anomalies for compromised accounts, report export volume for exfiltration, deployment monitoring for unauthorized code, permission change monitoring for escalation |
| Tabletop Exercises | ✅ Conduct tabletop exercises testing response procedures for platform-specific scenarios validating team readiness and procedure completeness |
| Response Automation | ✅ Real-time Transaction Security policies automatically block suspicious actions, require step-up authentication, or notify responders on high-confidence detection signals, providing automated containment that reduces response time |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Generic Response Plans Only | ⚠️ Having only generic incident response procedures without Salesforce-specific guidance, improvising platform-specific actions during active incidents |
| No Detection Configuration | ⚠️ Having no monitoring configured to detect platform-specific scenarios (bulk export, privilege escalation, unauthorized deployment), discovering incidents through user reports or audits |
| No Tabletop Exercises | ⚠️ Never conducting tabletop exercises validating team familiarity with platform-specific response procedures, discovering knowledge gaps during actual incidents |
| No Automation | ⚠️ Fully manual response for all incidents with no real-time Transaction Security policies for automated containment, introducing delays when immediate blocking could reduce damage |
| Missing Response Playbooks | ⚠️ Lacking documented response playbooks for common scenarios (compromised user account, data exfiltration, unauthorized code deployment) |
| No Breach Notification Procedures | ⚠️ Having no procedures for regulatory breach notification requirements, scrambling to understand obligations during actual incidents when timelines are tight |
Patterns
| Where to look | What good looks like |
|---|---|
| Out-of-Band Communication | ✅ Establish communication channels that do not depend on potentially compromised systems (phone trees, external messaging platforms) for coordinating response |
| Escalation Procedures | ✅ Document escalation paths from initial detection through executive notification with contact information and escalation criteria |
| Notification Templates | ✅ Maintain pre-drafted notification templates for common incident types enabling rapid communication to stakeholders, customers, and regulators |
| Stakeholder Identification | ✅ Identify stakeholders requiring notification for different incident types including executives, legal, compliance, affected users, and regulatory authorities |
| Regulatory Notification Timelines | ✅ Document each jurisdiction’s regulatory notification requirements, confirmed against the governing regulation for your deployment, with procedures ensuring timely compliance |
| Internal Coordination | ✅ Design incident coordination procedures specifying roles and responsibilities (incident commander, technical lead, communications lead, legal liaison) |
| External Coordination | ✅ Establish procedures for coordinating with Salesforce Support, external forensic investigators, legal counsel, and law enforcement as appropriate |
| Status Communication | ✅ Implement regular status communication cadence during active incidents keeping stakeholders informed without overwhelming investigation team |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| No Out-of-Band Communication | ⚠️ Depending entirely on potentially compromised email or internal messaging for incident coordination, losing communication capability when primary channels are compromised |
| Unclear Escalation Paths | ⚠️ Having no documented escalation procedures, causing delays and confusion about who should be notified and when |
| No Notification Templates | ⚠️ Improvising incident notifications during active incidents under time pressure, resulting in incomplete or incorrect communications |
| Missing Stakeholder Lists | ⚠️ Having no pre-identified stakeholder lists, discovering during incidents that critical parties (legal, compliance, executives) were not notified |
| No Regulatory Timeline Awareness | ⚠️ Being unaware of the regulatory notification timelines that apply to your deployment during incident response, missing deadlines and creating additional compliance violations |
| Ambiguous Roles | ⚠️ Having unclear incident response roles and responsibilities, causing confusion about who has authority to make containment decisions |
| No External Coordination Procedures | ⚠️ Having no established relationships or procedures for engaging Salesforce Support, forensic investigators, or law enforcement when needed |
Patterns
| Where to look | What good looks like |
|---|---|
| Blameless Culture | ✅ Focus post-incident reviews on system improvements rather than individual blame, enabling organizational learning without fear of repercussions |
| Incident Timeline Documentation | ✅ Document complete incident timeline including initial compromise, detection, response actions, and resolution with precise timestamps |
| Control Failure Analysis | ✅ Identify which preventive and detective controls failed to prevent or detect the incident, informing architectural improvements |
| Architectural Weakness Documentation | ✅ Document architectural weaknesses revealed by the incident in Architecture Decision Records (ADRs) capturing security changes |
| Remediation Prioritization | ✅ Prioritize remediation based on risk reduction potential, addressing root causes rather than symptoms |
| Lesson Sharing | ✅ Share lessons learned across teams and business units preventing similar incidents in other parts of the organization |
| Detection Rule Updates | ✅ Update detection rules and behavioral baselines based on attack techniques observed during investigation |
| Response Procedure Refinement | ✅ Update incident response procedures based on what worked and what didn’t during actual incident response |
| Metrics Tracking | ✅ Track mean time to detect (MTTD), mean time to respond (MTTR), and scope of impact over time identifying trends requiring architectural investment |
| Review Cadence | ✅ Conduct a timely post-incident review while details remain fresh and momentum for improvement exists |
Anti-Patterns
| Where to look | What bad looks like |
|---|---|
| Skipped Reviews | ⚠️ Skipping post-incident reviews, moving directly to next incident without capturing lessons learned or architectural improvements |
| Blame Culture | ⚠️ Focusing post-incident reviews on individual blame rather than system improvements, causing incident concealment and preventing organizational learning |
| No Architectural Documentation | ⚠️ Failing to document security architecture changes driven by incident learnings in Architecture Decision Records (ADRs) |
| Symptom Treatment | ⚠️ Addressing only incident symptoms (compromised password) without fixing root causes (lack of MFA), allowing similar attacks to succeed |
| No Lesson Sharing | ⚠️ Keeping incident learnings within response team rather than sharing across organization, allowing similar incidents in other business units |
| No Detection Updates | ⚠️ Failing to update detection rules or behavioral baselines based on observed attack techniques, remaining vulnerable to similar attacks |
| No Metrics Tracking | ⚠️ Not tracking mean time to detect (MTTD), mean time to respond (MTTR), or scope trends, missing opportunities to identify systemic weaknesses |
| Delayed Reviews | ⚠️ Conducting post-incident reviews months after incident resolution when details are forgotten and improvement momentum is lost |
| No Remediation Follow-Through | ⚠️ Identifying improvements during post-incident review but failing to implement them, allowing similar incidents to recur |