Between approximately 15:00 UTC and 17:00 UTC on September 23, jobs took longer than usual to start, and some jobs were cancelled mid-run. Overall demand for our runners exceeded our dedicated capacity, and larger runners (16 and 30 vCPUs) were hit hardest, since they're the hardest to place when capacity is tight.
To absorb the extra demand, we ran some jobs on external compute resources. Some of those machines were reclaimed mid-run, which cancelled their jobs. Handling the lost machines slowed down our control plane, so jobs on our own servers also took longer to start. The more demand grew, the more interruptions we saw, and the further provisioning fell behind.
We stopped using reclaimable external compute resources and added more dedicated capacity. We're also changing how our control plane handles unreachable machines, so problems with external compute resources can't slow down provisioning on our own servers. This incident is resolved.