Our investigation determined that an internal service responsible for monitoring and resource state management became unresponsive, which contributed to the issues reported during the incident, including instances remaining in a Resuming state longer than expected, resource state discrepancies, and under-capacity alerts.
Our technical teams performed recovery actions, including restarting the affected service, and have since observed stable service operation. Following validation and monitoring, we have confirmed that affected services have returned to normal operation.
This incident is now considered resolved. A post-incident review is ongoing, and we will continue to evaluate opportunities for improvement.
Posted Aug 25, 2026 - 22:57 PDT
Monitoring
We are continuing to observe stable service behavior across the affected areas, and the previously reported symptoms are no longer being observed. Our technical teams are closely monitoring the environment and validating service stability across affected customer workloads.We will continue to monitor and provide a further update once monitoring activities are complete or if additional information becomes available.
Posted Aug 25, 2026 - 21:41 PDT
Update
Our technical teams are observing early signs of recovery and are actively monitoring service behavior while validating the impact across affected customer workloads.
Investigation and recovery efforts remain ongoing. We will continue to provide updates as we learn more and confirm service stability.
Posted Aug 25, 2026 - 21:23 PDT
Investigating
Incident Description: We are investigating an issue affecting portions of the Spot platform. Customers may experience instances remaining in a Resuming state longer than expected, under-capacity alerts, and discrepancies between resource status information displayed in Spot and AWS. In some cases, customers may observe differences between the status shown in the Spot console and the actual status reported by AWS for affected resources. Our technical teams are actively investigating the scope and impact of the issue.
Priority: P2
Restoration Activity: Our technical teams are actively investigating the issue and assessing its impact across affected services and customer workloads. Current efforts are focused on identifying the underlying cause, validating the full scope of customer impact, and restoring normal service behavior. Further updates will be provided as more information becomes available.