Salesforce Outage Chaos: System Failures Spread Across Multiple Regions
Salesforce's September 16 outage lasted for 7 hours and 36 minutes, affecting multiple regions and causing errors, delays, and unavailability of services. The incident started at 07:50 UTC and was caused by requests getting stuck while waiting on an internal login service, which used up available server resources.
Salesforce tried to recover the system with rolling restarts, blocking API endpoints, and pushing fixes across regions. However, some environments still needed manual restarts after interactive access returned. The outage also disrupted downstream workflows, making it difficult for customers to create support cases through Salesforce Help.
Twilio's carrier incidents highlighted the importance of separate monitoring for message acceptance, delivery, and delivery confirmation. Akamai and GitHub added two lessons: keeping an alternative support route and testing journeys instead of relying on a single healthy metric.