Ubicloud - Notice history

Ubicloud Runners - Operational

GitHub Actions - Operational

GitHub API Requests - Operational

PostgreSQL - Operational

Compute - Operational

Networking - Operational

Kubernetes - Operational

GitHub Webhooks - Operational

Ubicloud Cache - Operational

Web Console - Operational

Compute - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

PostgreSQL - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Networking - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Kubernetes - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Web Console - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Notice history

View current status

Oct 2026

Actions Job Delays
ResolvedDegraded performance3 hours 9 minutes
  • Resolved
    UTC
    Resolved

    This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

  • Update
    UTC
    Update

    GitHub Actions experienced degraded performance for some hosted runners due to throttling within an upstream Azure dependency. Service capacity has recovered, and we are continuing to monitor while working with Azure on the underlying condition.

  • Update
    UTC
    Update

    We are currently applying a mitigation and anticipate recovery within thirty minutes.

  • Update
    UTC
    Update

    We have identified an issue with our upstream provider which is causing Actions requests to 429 which is creating the delays. We have escalated to the owning team and are investigating how to mitigate the 429s.

  • Update
    UTC
    Update

    We are seeing a reoccurrence in run-start delays, and are continuing to investigate the issue. Customers will potentially experience delays of up to ten minutes.

  • Investigating
    UTC
    Investigating

    Actions is experiencing degraded performance. We are continuing to investigate.

  • Monitoring
    UTC
    Monitoring

    The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    Run-start delays on Ubuntu runners have been resolved. We are investigating run-start delays on Windows runners and will share more information as it becomes available.

  • Update
    UTC
    Update

    We have identified the cause of increased Actions run start delays on Ubuntu runners and are actively deploying a fix. Customers may continue to experience intermittent delays while the mitigation rolls out and service metrics return to normal.

  • Investigating
    UTC
    Investigating

    We are investigating reports of degraded performance for Actions

Sep 2026

Increased provisioning times for large runner sizes (16 and 30 vCPUs)
ResolvedDegraded performance1 hour 45 minutes
  • Resolved
    UTC
    Resolved

    Between approximately 15:00 UTC and 17:00 UTC on September 23, jobs took longer than usual to start, and some jobs were cancelled mid-run. Overall demand for our runners exceeded our dedicated capacity, and larger runners (16 and 30 vCPUs) were hit hardest, since they're the hardest to place when capacity is tight. To absorb the extra demand, we ran some jobs on external compute resources. Some of those machines were reclaimed mid-run, which cancelled their jobs. Handling the lost machines slowed down our control plane, so jobs on our own servers also took longer to start. The more demand grew, the more interruptions we saw, and the further provisioning fell behind. We stopped using reclaimable external compute resources and added more dedicated capacity. We're also changing how our control plane handles unreachable machines, so problems with external compute resources can't slow down provisioning on our own servers. This incident is resolved.

  • Investigating
    UTC
    Investigating

    We're seeing higher than usual demand for our runners. This is causing longer than usual provisioning times for larger runners (16 and 30 vCPUs). We're working to increase our capacity as quickly as possible and mitigate the issue.

Increased provisioning times for large runner sizes (16 and 30 vCPUs)
ResolvedDegraded performance1 hour 45 minutes
  • Postmortem
    UTC
    Postmortem

    Between approximately 15:00 UTC and 17:00 UTC on September 23, jobs took longer than usual to start, and some jobs were cancelled mid-run. Overall demand for our runners exceeded our dedicated capacity, and larger runners (16 and 30 vCPUs) were hit hardest, since they're the hardest to place when capacity is tight.

    To absorb the extra demand, we ran some jobs on external compute resources. Some of those machines were reclaimed mid-run, which cancelled their jobs. Handling the lost machines slowed down our control plane, so jobs on our own servers also took longer to start. The more demand grew, the more interruptions we saw, and the further provisioning fell behind.

    We stopped using reclaimable external compute resources and added more dedicated capacity. We're also changing how our control plane handles unreachable machines, so problems with external compute resources can't slow down provisioning on our own servers. This incident is resolved.

  • Resolved
    UTC
    Resolved

    All systems are operating normally.

  • Investigating
    UTC
    Investigating

    We're seeing higher than usual demand for our runners. This is causing longer than usual provisioning times for larger runners (16 and 30 vCPUs). We're working to increase our capacity as quickly as possible and mitigate the issue.

Incident across several services
ResolvedDegraded performance18 hours 44 minutes
  • Resolved
    UTC
    Resolved

    Starting at 07:57 UTC on September 23, GitHub experienced elevated 500 and 404 responses across several application pages. This caused failures when installing GitHub Apps, creating organizations, and making some organization membership changes. Customers also experienced delayed label updates and stale search results in Projects.

    The infrastructure failure was isolated to our Azure Central US region. The API errors were mitigated by 10:58 UTC on September 23. Projects' processing continued to recover while an accumulated backlog was drained, and full service was restored at 04:55 UTC on September 24.

    The incident was caused by a failed planned maintenance operation on a primary database. Automated recovery initiated an emergency database failover, after which several replicas in the affected region were unable to resume replication correctly. This reduced available database capacity and caused the API errors and downstream Projects processing delays.

    We have mitigated the immediate failure mode. We are also improving maintenance safety checks, database failover handling, post-failover replica validation, and downstream processing resilience to reduce the likelihood and impact of similar incidents.

  • Monitoring
    UTC
    Monitoring

    The degradation has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    We are continuing to process the backlog of issue label updates for Projects. Label changes may still be delayed. All other services are operating normally.

  • Update
    UTC
    Update

    We are continuing to process the backlog of issue label updates for Projects. Users may still see delays before label changes are reflected in Projects. All other services are operating normally.

  • Update
    UTC
    Update

    We've deployed a change intended to accelerate processing of the backlog of issue label updates in Projects. A sizable backlog still remains and we continue working through it. All other services are operating normally. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Updates to issue labels may be delayed in being reflected in Projects. We are continuing to deploy a change that will accelerate processing of the backlog of label updates. All other services are available. We will provide another update within the next hour.

  • Update
    UTC
    Update

    We are preparing to deploy a change that will mitigate the impact.

  • Update
    UTC
    Update

    Continuing to investigate the lag that may be experienced in issue labels being accurately reflected in Projects. We are working on alternate solutions to process the backlog of label updates.

  • Update
    UTC
    Update

    We will post another update in approximately one hour to share our progress.

  • Update
    UTC
    Update

    Updates to issue labels may be delayed in being reflected in Projects by about ~10 minutes. We have added some capacity to work through the backlog more quickly, but it'll likely be a few hours to complete processing the full backlog of messages. All other services are available.

  • Update
    UTC
    Update

    Users may experience stale Project search results. We are working to increase indexing speed. All other services are available.

  • Update
    UTC
    Update

    We are seeing recovery for Projects. Users may experience stale search results for Projects while indexing catches up.

  • Update
    UTC
    Update

    The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    Database replicas have been restored. Org creation and the GitHub API are no longer degraded.

  • Update
    UTC
    Update

    Database replicas have detached. We're working to restore the database replicas. Users may experience issues beyond creating organizations and a degraded experience with the GitHub API and Projects.

  • Investigating
    UTC
    Investigating

    We are investigating reports of degraded performance for API Requests

Aug 2026

Incident with Actions and Pull Requests
ResolvedDegraded performance1 hour 30 minutes
  • Resolved
    UTC
    Resolved

    On August 26, 2026, from 21:55 UTC to 23:58 UTC, 2.6% of workflow runs triggered by pull request events were delayed, with the impact rising as high as 25% at its peak. Some users also experienced delays in pull request merge-commit generation, mergeability information, and merge-button availability. Actions and Pull Requests fully recovered by 23:58 UTC; the incident was resolved at 00:26 UTC after normal operation was confirmed.

    Background jobs that process pull request updates and generate merge commits were impacted by timeouts reaching a single partition of git data. This resulted in a backlog in pull request merge-commit processing, delaying pull request-triggered GitHub Actions workflows and some mergeability information.

    We reduced workload, shifted traffic away from affected infrastructure, and restored the affected service component to a healthy state. Together, these actions helped drain the backlog and restore normal operations.

    We are working to improve resource saturation detection and to eliminate customer impact in this scenario by isolating impact, placing better bounds on retries, and strengthening backpressure to make our systems more resilient under load.

  • Monitoring
    UTC
    Monitoring

    The degradation affecting Actions and Pull Requests has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    We confirmed full recovery beginning at 23:58 UTC. Actions workflow runs and pull request merges are operating normally. We will now resolve the incident while continuing to monitor service health.

  • Update
    UTC
    Update

    We've applied mitigations and are seeing recovery in Actions workflow runs and blocked pull request merges. We're continuing to monitor for sustained health of merge commit creates before resolving.

  • Update
    UTC
    Update

    We are investigating elevated delays and timeouts affecting Actions workflow runs triggered by pull request events. 20% of actions runs have delayed starts of more than 5 minutes and up to 4% of runs failed to trigger. We are actively working on mitigation and will provide updates as we learn more.

  • Investigating
    UTC
    Investigating

    We are investigating reports of degraded performance for Actions and Pull Requests

Incident with Actions
ResolvedMajor outage2 hours 49 minutes
  • Resolved
    UTC
    Resolved

    On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system.

    At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC.

    3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs.

    Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27.

    Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs.

    Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.

  • Update
    UTC
    Update

    All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.

  • Monitoring
    UTC
    Monitoring

    The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.

  • Update
    UTC
    Update

    We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.

  • Update
    UTC
    Update

    Pages is operating normally.

  • Update
    UTC
    Update

    We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

  • Update
    UTC
    Update

    primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

  • Update
    UTC
    Update

    We've identified an issue with a database primary and are failing over to a replica immediately

  • Update
    UTC
    Update

    Pages is experiencing degraded performance. We are continuing to investigate.

  • Investigating
    UTC
    Investigating

    We are investigating reports of degraded availability for Actions

Actions delays in starting runs
ResolvedDegraded performance37 minutes
  • Resolved
    UTC
    Resolved

    On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright.

    The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC.

    To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.

  • Monitoring
    UTC
    Monitoring

    The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.

  • Investigating
    UTC
    Investigating

    We are investigating reports of degraded performance for Actions

Incident with GitHub.com
ResolvedMajor outage7 hours 35 minutes
  • Resolved
    UTC
    Resolved

    On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02.

    Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.

    The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.

    The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed.

    Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery.

    Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints.

    To prevent recurrence, our follow-up actions include:

    - Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity.

    - Auditing Istio request, concurrency, and scaling limits across affected services.

    - Reviewing retry limits and backoff behavior across gateways and clients.

    - Addressing the VS Code retry behavior that amplified Copilot token traffic.

    - Improving load-balancer capacity monitoring and regional failover safeguards.

  • Update
    UTC
    Update

    We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.

  • Update
    UTC
    Update

    Issues is operating normally.

  • Update
    UTC
    Update

    We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.

  • Update
    UTC
    Update

    We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.

  • Update
    UTC
    Update

    API Requests is operating normally.

  • Update
    UTC
    Update

    API Requests is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.

  • Update
    UTC
    Update

    Issues is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.

  • Update
    UTC
    Update

    Git Operations is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

  • Update
    UTC
    Update

    We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.

  • Update
    UTC
    Update

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.

  • Update
    UTC
    Update

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.

  • Update
    UTC
    Update

    Webhooks is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    Git Operations is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    Pages is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    API Requests is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    Webhooks is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement.

  • Update
    UTC
    Update

    Actions is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    Pull Requests is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    Issues is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    Pull Requests is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    Copilot is experiencing degraded availability. We are continuing to investigate.

  • Update
    UTC
    Update

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information.

  • Update
    UTC
    Update

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate.

  • Update
    UTC
    Update

    Pull Requests is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    Issues is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available

  • Update
    UTC
    Update

    Webhooks is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    Actions is experiencing degraded performance. We are continuing to investigate.

  • Update
    UTC
    Update

    API Requests is experiencing degraded performance. We are continuing to investigate.

  • Investigating
    UTC
    Investigating

    We are investigating reports of impacted performance for some GitHub services.

Previous

Aug 2026 to Oct 2026

Next