Buildkite logo and current status indicator

Buildkite Incident History

Operational

buildkite.com 12 components

checked Sep 11, 2026 4:18 AM UTC · Buildkite's official status page

Buildkite is up and running.

Buildkite is currently operational with all systems functioning normally.

All components operational

Email only. No password. No card.

Be the first to know

Statusfield watches Buildkite alongside the rest of your stack and tells you when its status changes. Email, Slack, Microsoft Teams, PagerDuty, and webhooks.

Email only. No password. Free.

Buildkite incident history

11 reports published by Buildkite in the last 30 days

When Buildkite breaks, it is typically resolved in 1h 6m — median across 22 resolved incidents over 90 days.

September 10, 2026

Buildkite service disruption

Last update
  1. resolvedSep 10 · 01:54 UTC

    We have seen full recovery for customers since 20:28 UTC. We experienced an autoscaling feedback loop which increased the number of connections to our redis cluster above its ability to respond. This had widespread impact for all of our customers with Web UI, Agent API, REST API and job queue impact between 18:43-19:21 UTC, and again between 20:02-20:28 UTC. A full post incident review will be available later this week.

Reported by Buildkite on their status page.

September 4, 2026

Hosted Agents - Increased latency and error rates

Last update
  1. resolvedSep 4 · 02:39 UTC

    Jobs running on Hosted Agents are now being dispatched promptly.

Reported by Buildkite on their status page.

September 3, 2026

Scheduling delay for hosted agents

Last update
  1. resolvedSep 3 · 07:06 UTC

    Jobs running on Hosted Agents are now being dispatched promptly.

Reported by Buildkite on their status page.

September 2, 2026

Delays in Hosted Agents job dispatch

Last update
  1. resolvedSep 2 · 15:51 UTC

    Job dispatch for Hosted Agents is no longer experiencing delays.

Reported by Buildkite on their status page.

September 1, 2026

Delivery issues with email notifications

Last update
  1. resolvedSep 1 · 16:52 UTC

    Our outbound email queue has finished processing all previously failed deliveries. Email notifications are processing successfully.

Reported by Buildkite on their status page.

August 27, 2026

Buildkite service disruption

Last update
  1. postmortemAug 27 · 23:53 UTC

    ## Service Impact Customers experienced a site-wide disruption affecting the Buildkite web interface, REST API, Agent API, job queue, and notifications. The disruption began at 22:47 UTC and continued until 23:16 UTC. During this period, all customers were unable to create new builds via the UI or REST API, and job dispatch was delayed. We continued to receive webhooks from source control providers, but processing of those webhooks was delayed. New build processing resumed at 23:10 UTC and the accumulated backlog of work was processed by 23:16 UTC. Some builds that were already in progress at the beginning of the impact period remained temporarily stuck until services recovered. Some customers also experienced errors in the web UI, and on receiving webhooks from SCM providers. This impacted <1.5% of all requests during the impact period. ## Incident Summary **Background** Buildkite services running in one of our production Kubernetes clusters depend on an internal DNS service CoreDNS to locate databases, queues, and other application components. Over the last four months, we have been migrating our production workloads from AWS ECS to this AWS EKS cluster. The cluster has been steadily increasing in size during the course of this migration. Additionally, Buildkite recently moved time-sensitive notification jobs from a general-purpose pool of background workers into a new low-latency worker pool. To ensure sufficient capacity for both pools, we initially configured each with the same high maxReplica count as the original shared pool, with the intention of reviewing and adjusting the limits for both pools downwards at a later date. **Trigger: A sudden increase in demand on CoreDNS** At 22:44 UTC, an application deploy created a surge in application Pod volume, which caused our EKS cluster to scale out. This surge consumed the available headroom on already-deployed Nodes, which limited applications’ capacity to autoscale promptly. Some background workers \(including the aforementioned notification workers\) were also attempting to scale out at this time. The headroom shortage and subsequent cluster autoscaling delayed provisioning of the compute requested by those services. Since there had been no change in the metrics that triggered the services to scale up, those services requested even more Pods. Normally the impact of such runaway autoscaling would be limited by the services’ configured maximums. However, as mentioned previously the maximums for these services had been set higher than usual. The application deployment and runaway autoscaling combined to trigger an unusually high rate of change to applications, network endpoints, and cluster nodes. **Why CoreDNS failed** The cluster's CoreDNS service was running at a fixed size and did not automatically scale with the size or rate of change of the cluster. During post-incident analysis we discovered some bugs and gaps in our monitoring of CoreDNS, particularly around query volume and duration. * We found a defect in a key monitoring query which had masked an upward trend in query duration, correlated with increasing cluster size. * We also found that DNS query volume was not sufficiently monitored, and had been trending upwards as we migrated more workloads into EKS. These defects masked the upward trend in latency on the CoreDNS side; the client-side latency increase was offset by recent performance gains in our applications, and so did not catch our attention. Hence, we had not correctly prioritised our planned implementation of autoscaling for CoreDNS. **What happened when CoreDNS failed** During the impact period CoreDNS’s completed query rate remained stable, but query processing time increased from under 1ms to approximately 780ms. Pending requests accumulated in memory, until all three original CoreDNS pods exceeded their allowed memory limits and were restarted by Kubernetes. Continued demand and DNS retries prevented the service from recovering after restarting. CoreDNS unavailability caused failures across APIs, job dispatch, and notifications for all customers. Retries and delayed work increased the load during recovery. As part of the migration from ECS to EKS, database queries for a subset of customers were routed to PgBouncer instances running in the EKS cluster. During the impact period, these customers’ requests to webhooks and the Web UI received error responses. It was less than 1.5% of all requests during this incident that returned errors. **How we responded** We paused further application deployments, deployed more CoreDNS service replicas, raised the memory available to each CoreDNS replica, and expanded the node pool available to run them. The new set of CoreDNS Pods came into service by 23:12 UTC. After that, DNS errors fell rapidly. Customer-facing services processed the accumulated backlog and recovered fully by 23:16 UTC. ## Changes we're making * **Increase immediate CoreDNS headroom.** We increased CoreDNS from three to twelve pods, raised each pod’s memory limit from 512 MiB to 2 GiB, and expanded the system node pool available to run them. * **Enable automatic CoreDNS capacity scaling.** We will enable EKS-managed CoreDNS autoscaling, with a tested minimum replica count and sufficient capacity to distribute those replicas across nodes and availability zones. * **Detect both rapid degradation and declining headroom.** We will add direct alerts for CoreDNS query duration, query volume, goroutine growth, memory pressure, OOM restarts, and available replicas. We will also monitor longer-term trends as part of capacity planning. * **Apply the same standard to other cluster-critical services.** We will identify and remediate critical services that lack tested capacity, direct alerting, failure-domain distribution, and either safe autoscaling or documented static headroom.

Reported by Buildkite on their status page.

Test Engine ingestion delayed, catching up

Last update
  1. resolvedAug 27 · 11:47 UTC

    Processing of uploaded test execution data has caught up and is back to normal.

Reported by Buildkite on their status page.

August 26, 2026

Issues identified with buildkite-agent registration

Last update
  1. resolvedAug 26 · 03:13 UTC

    Mitigation strategies worked as expected and now agent registration is back to normal. All systems are stable now.

Reported by Buildkite on their status page.

August 25, 2026

Unexpected agent disconnection for some customers

Last update
  1. resolvedAug 25 · 08:16 UTC

    This incident has been resolved.

Reported by Buildkite on their status page.

August 17, 2026

Buildkite service disruption

Last update
  1. resolvedAug 17 · 02:43 UTC

    This incident has been resolved.

Reported by Buildkite on their status page.

August 13, 2026

Slow web UI

Last update
  1. resolvedAug 13 · 16:40 UTC

    We think the impact from the issue is over.

Reported by Buildkite on their status page.