Ubicloud - Increased provisioning times for large runner sizes (16 and 30 vCPUs) – Incident details

Increased provisioning times for large runner sizes (16 and 30 vCPUs)

Resolved
Degraded performance
Started 9 days agoLasted 1 hour 45 minutes

Affected

GitHub Actions Service

Degraded performance from 12:14 PM to 1:59 PM

Ubicloud Runners

Degraded performance from 12:14 PM to 1:59 PM

Updates
  • Postmortem
    UTC
    Postmortem

    Between approximately 15:00 UTC and 17:00 UTC on September 23, jobs took longer than usual to start, and some jobs were cancelled mid-run. Overall demand for our runners exceeded our dedicated capacity, and larger runners (16 and 30 vCPUs) were hit hardest, since they're the hardest to place when capacity is tight.

    To absorb the extra demand, we ran some jobs on external compute resources. Some of those machines were reclaimed mid-run, which cancelled their jobs. Handling the lost machines slowed down our control plane, so jobs on our own servers also took longer to start. The more demand grew, the more interruptions we saw, and the further provisioning fell behind.

    We stopped using reclaimable external compute resources and added more dedicated capacity. We're also changing how our control plane handles unreachable machines, so problems with external compute resources can't slow down provisioning on our own servers. This incident is resolved.

  • Resolved
    UTC
    Resolved

    All systems are operating normally.

  • Investigating
    UTC
    Investigating

    We're seeing higher than usual demand for our runners. This is causing longer than usual provisioning times for larger runners (16 and 30 vCPUs). We're working to increase our capacity as quickly as possible and mitigate the issue.