Identified increased agent join latencies in eu-central
Last updatepostmortemSep 8 · 21:59 UTC
Root Cause
A configuration change introduced a formatting error in the startup settings for the server pools that run hosted agents in our eu-central region. Newly provisioned servers failed to start, so the region could not add capacity. Agents that were already running were unaffected, but new agent deployments, version rollouts, and automatic scale-ups were delayed or stuck pending until the change was reverted.Timeline (UTC)
2026-09-04 23:56- Configuration change deployed. The error only affected newly provisioned servers, so there was no immediate impact.
2026-09-05 18:07- New servers began failing to start. Impact begins, intermittent at first.
2026-09-05 22:23- Our monitoring alerted us to agent deployments stuck pending and we began investigating.
2026-09-06 00:50- We identified the malformed startup configuration and reverted the change.
2026-09-06 01:17- New capacity provisioned successfully, all pending deployments recovered, and we validated the fix.Scope of Impact
Limited to hosted agents in our eu-central region. During the impact window, new agent deployments and rollouts were delayed or stuck pending, and automatic scale-up was blocked. Agents already running continued to serve sessions normally, and we estimate that, at the peak, roughly 1% of agent instances in the region were affected. No other regions or products were affected.Mitigations and Follow-ups
- The faulty configuration change was reverted and provisioning has been stable since.
- We are adding alerting for new servers that fail to start via this failure mode, which today fails silently. This lets us detect this class of failure directly and validate future changes to server provisioning quickly.
We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.
Reported by LiveKit on their status page.
