Between 11:57 PM PDT on October 19 and 2:25 PM PDT on October 20, Spruce experienced a platform-wide disruption due to a major AWS outage and a concurrent Twilio issue.
There were three periods of major outage when most Spruce functionality was unavailable:
Outside these windows, the platform remained partially degraded.
The outage affected inbox access, secure messaging, calls, SMS, faxes, emails, video visits, and the public API. Services already running on healthy AWS instances continued to function intermittently, but scaling and communication components were heavily impacted.
The incident stemmed from widespread failures in AWS control-plane components (including EC2 orchestration, DynamoDB, Lambda, and ECR), preventing new ECS tasks from launching and disrupting core Spruce services. Twilio's concurrent outage compounded the impact, delaying or blocking call and message delivery.
In response, we've updated our on-call notification schedule for faster follow-the-sun escalations to our non-US engineers, and introduced process improvements and clearer communication practices to streamline incident response and keep users better informed. We're also evaluating backup system independence from AWS control-plane services and improvements to routing throughput and retry handling.
These steps aim to improve both resilience and response speed for future large-scale incidents.