Regulated Data Environment - Services degraded due to underlying cloud provider issue – Incident details

All systems operational

Services degraded due to underlying cloud provider issue

Resolved
Major outage
Started 10 months agoLasted about 21 hours

Affected

AWS Research & Engineering Studio

Major outage from 7:38 AM to 4:23 AM

AWS ParallelCluster

Partial outage from 7:38 AM to 11:22 PM, Degraded performance from 11:22 PM to 4:23 AM

Open OnDemand

Operational from 7:38 AM to 2:16 PM, Partial outage from 2:16 PM to 11:22 PM, Operational from 11:22 PM to 4:23 AM

Updates
  • Resolved
    UTC
    Resolved

    All systems are back to normal operational state.

    We will continue monitoring the situation tomorrow and complete some cleanup tasks.

    We will provide further updates if we discover any remaining issue.

  • Monitoring
    UTC
    Monitoring

    AWS has almost completely resolved the underlying infrastructure issues.

    We are still experiencing some issues with the Research Engineering studio.

    Parallel cluster should be working correctly.

    We will continue working to restore full functionality and we will keep providing updates.

  • Identified
    UTC
    Identified

    Some of the underlying AWS services we leverage in our computational offerings are still experiencing issues.

    We keep monitoring AWS status updates and assess the impact on internal services.

  • Resolved
    UTC
    Resolved

    The cloud provider has reported that the underlying root cause of the problems have been corrected.
    This incident has been resolved.

  • Identified
    UTC
    Identified

    Several cloud services in the us-east-1 region in AWS are currently affected by multiple failures. AWS is investigating.

    Our computational and data transfer services are affected and either failing to operate or operating with degraded functionalities. We are monitoring and waiting for the cloud provider to resolve the underlying issues.