Description: Software Vulnerability Research (SVR) - Service Disruption
Timeframe: July 28, 2026, 2:00 PM PDT to July 29, 2026, 2:53 AM PDT
Incident Summary
On Tuesday, 28 July 2026, at 2:00 PM PDT, the Flexera Software Vulnerability Research(SVR) production environment experienced a service disruption that affected API availability and application functionality. The incident occurred when the primary database instance supporting the SVR platform became unavailable. Existing database connections were unexpectedly terminated, and application servers were unable to establish new connections. As a result, customers were unable to reliably access SVR services and APIs during the incident.
During the investigation, our teams determined that the disruption was caused by an unplanned failover initiated by the cloud service provider. According to the provider’s event history, an infrastructure issue was detected on the primary database host, prompting the provider to automatically initiate the failover process.
Our technical teams immediately engaged the cloud service provider to investigate the incident and restore service. Following recovery and validation activities, all production services were successfully restored and returned to normal operation by 2:53 AM PDT on 29 July 2026.
Root Cause
The incident was triggered by an unplanned failover of the production database, initiated by the cloud service provider in response to an infrastructure-level issue affecting the primary database host.
During the failover, the primary database instance became temporarily unavailable. This caused active database connections to terminate abruptly, preventing application services from establishing new connections. As a result, API requests failed and service availability across the SVR platform was impacted.
Flexera has formally requested a detailed Root Cause Analysis (RCA) from the cloud service provider to determine the precise infrastructure condition that triggered the failover and to identify opportunities to prevent recurrence.
Contributing Factors
During the investigation, the following factors were identified as potentially contributing to the overall impact of the incident:
A significantly elevated volume of client connection attempts was observed originating from customer environments during the incident period.
Connection volumes exceeded normal operating levels, increasing load on backend services while database services were recovering.
Elevated connection activity may have amplified the impact of the database failover and increased recovery complexity.
Existing application retry and connection behaviours generated additional connection demand during database recovery activities.
Remediation Actions
Cloud provider engagement: Our teams immediately engaged the cloud service provider and worked directly with their on-call engineers throughout the recovery effort.
Database recovery: Executed a controlled database failover/reboot in coordination with the cloud service provider.
Restored connectivity: Re-established database availability and connectivity to the production environment.
Application recovery validation: Confirmed that application services successfully reconnected to backend database services.
Service validation: Verified recovery of APIs and critical SVR application functionality.
Post-recovery monitoring: Implemented enhanced monitoring of database health, application performance, and connection activity following restoration.
Future Preventative Measures
Continue monitoring customer connection volumes and backend database connection utilisation.
Review and optimise database connection pooling, connection management, retry logic, and failover handling to improve resilience during future infrastructure events.
Complete the cloud service provider support engagement and review the provider's detailed Root Cause Analysis once available.
Conduct an internal post-incident review and implement additional preventive measures identified through the review process.