Pipelines are stuck in Prod1
Last updatepostmortemSep 10 · 07:21 UTC
# Summary On September 3, 2026, between approximately 2:54 AM and 3:44 AM PDT, customers running pipelines on Prod1 experienced delays in pipeline execution graph rendering and execution status updates. Pipeline execution itself was not affected — pipelines continued to run and complete — but the visual graph and status information in the Harness UI lagged behind actual execution progress. Harness engineering identified the cause, added processing capacity, and restored normal graph and status updates. Extended monitoring confirmed full recovery the following day. # Root Cause A planned database failover activity in the Prod1 region temporarily increased network latency between the pipeline execution service and its database while traffic was briefly served cross-region. This reduced the rate at which the service could process pipeline orchestration events, and a processing backlog began to build. Pipeline execution graphs and status indicators in the UI depend on these orchestration events being processed in near real time. As the backlog grew, graph rendering and status updates fell increasingly behind actual execution progress, and some graphs had to be rebuilt rather than served from cache. Pipeline execution itself does not depend on this same processing path, so pipelines continued to run and complete throughout the incident. # Impact * Affected users: Customers with pipeline executions running on Prod1 during the incident window. * Symptom: Delayed pipeline execution graph rendering and delayed execution status updates in the Harness UI. * Not reported as impacted: Pipeline execution itself — pipelines continued to run and complete. * Data: No data loss, corruption, or exposure was identified. * Suggested workaround: None required from customers. The issue was resolved entirely by Harness engineering. * Duration of impact: Approximately 45–50 minutes of degraded graph and status visibility, from 2:54 AM to 3:44 AM PDT on September 3, 2026. Harness continued monitoring after the backlog cleared and confirmed full recovery. # Remediation ## Immediate Harness engineering added processing capacity to the affected service and restarted it, which cleared the event-processing backlog and restored normal pipeline execution graph and status updates. # Action Items | **Action Item** | **Objective** | | --- | --- | | Queue based alarms | We are raising priority for queue based alarms so we can take faster remediation actions \(increase capacity\) |
Reported by Harness on their status page.
