LiveKit logo and current status indicator

LiveKit Incident History

Operational

livekit.io 16 components

checked Sep 15, 2026 1:11 AM UTC · LiveKit's official status page

LiveKit is up and running.

LiveKit is currently operational with all systems functioning normally.

All components operational

Email only. No password. No card.

Be the first to know

Statusfield watches LiveKit alongside the rest of your stack and tells you when its status changes. Email, Slack, Microsoft Teams, PagerDuty, and webhooks.

Email only. No password. Free.

LiveKit incident history

13 reports published by LiveKit in the last 30 days · 2 still open

When LiveKit breaks, it is typically resolved in 1h 35m — median across 14 resolved incidents over 90 days.

September 5, 2026

Identified increased agent join latencies in eu-central

Last update
  1. postmortemSep 8 · 21:59 UTC

    Root Cause
    A configuration change introduced a formatting error in the startup settings for the server pools that run hosted agents in our eu-central region. Newly provisioned servers failed to start, so the region could not add capacity. Agents that were already running were unaffected, but new agent deployments, version rollouts, and automatic scale-ups were delayed or stuck pending until the change was reverted.

    Timeline (UTC)
    2026-09-04 23:56 - Configuration change deployed. The error only affected newly provisioned servers, so there was no immediate impact.
    2026-09-05 18:07 - New servers began failing to start. Impact begins, intermittent at first.
    2026-09-05 22:23 - Our monitoring alerted us to agent deployments stuck pending and we began investigating.
    2026-09-06 00:50 - We identified the malformed startup configuration and reverted the change.
    2026-09-06 01:17 - New capacity provisioned successfully, all pending deployments recovered, and we validated the fix.

    Scope of Impact
    Limited to hosted agents in our eu-central region. During the impact window, new agent deployments and rollouts were delayed or stuck pending, and automatic scale-up was blocked. Agents already running continued to serve sessions normally, and we estimate that, at the peak, roughly 1% of agent instances in the region were affected. No other regions or products were affected.

    Mitigations and Follow-ups

    • The faulty configuration change was reverted and provisioning has been stable since.
    • We are adding alerting for new servers that fail to start via this failure mode, which today fails silently. This lets us detect this class of failure directly and validate future changes to server provisioning quickly.

    We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

Reported by LiveKit on their status page.

September 4, 2026

Some sessions incorrectly shown as active

Last update
  1. resolvedSep 4 · 10:00 UTC

    A subset of room sessions that ended around 4 September 09:57 UTC were not recorded with an end time. These sessions continue to appear as active in the Cloud dashboard and in session records, and no room_ended end time is reflected for them.

    This was a reporting issue only. Rooms ended normally for participants, live traffic was not affected, and there is no impact on usage or billing.

    The underlying cause has been identified and fixed. Affected sessions from this window may continue to display as active; they can be safely disregarded.

Reported by LiveKit on their status page.

September 2, 2026

Investigating reports of one-way audio issues in LiveKit Phone Numbers

Last update
  1. resolvedSep 2 · 21:45 UTC

    This incident has been resolved. We will follow up with a detailed postmortem as soon as possible.

Reported by LiveKit on their status page.

Twilio Voice outage

Ongoing · last update
  1. monitoringSep 2 · 08:48 UTC

    Since ~06:30 UTC, outbound calls routed through Twilio are failing at an elevated rate. Twilio is reporting degraded voice connectivity (https://status.twilio.com/), affecting North America, Latin America, Europe, and the Middle East & Africa.

Reported by LiveKit on their status page.

LiveKit dashboard failing to load

Last update
  1. resolvedSep 2 · 11:57 UTC

    This incident has been resolved and dashboard load times returned to normal by 10:55 UTC. As a note, our beta agent simulations feature was also impacted from approximately 06:07 - 07:52 UTC.

Reported by LiveKit on their status page.

August 26, 2026

Elevated inference errors — US East

Last update
  1. resolvedAug 26 · 12:22 UTC

    Between 08:47 and 08:54 UTC, the service that routes LLM, STT and TTS requests through LiveKit Inference in our US East region was unavailable, causing inference requests from agents in that region to fail or time out. Sessions depending on those responses may have degraded or ended prematurely. Replacement capacity came online at 08:54 UTC and request volumes returned fully to normal levels by 09:02 UTC. We have confirmed no further impact since then. The cause was a loss of capacity during automated infrastructure maintenance. We have identified fixes to prevent a recurrence.

Reported by LiveKit on their status page.

August 25, 2026

Investigating elevated API errors in India

Ongoing · last update
  1. investigatingAug 25 · 10:33 UTC

    Between approximately 09:00 and 10:05 UTC, we observed intermittent elevated error rates on participant management APIs in the Mumbai region. Error rates have since returned to normal, and other regions were not affected. We are continuing to investigate the root cause and are monitoring the region.

Reported by LiveKit on their status page.

August 24, 2026

Agent session analytics displaying incorrect values in EU Central and India regions

Last update
  1. resolvedAug 24 · 14:14 UTC

    Agent session analytics in the Cloud dashboard have recovered and current data is reporting correctly for all affected projects. Some gaps may remain in historical data from the affected window (August 21–24 UTC); these will be backfilled.

Reported by LiveKit on their status page.

August 23, 2026

Egress recordings terminated early in US East

Last update
  1. resolvedAug 23 · 21:00 UTC

    From 09:18 UTC on 21 August to 21:00 UTC on 23 August, 0.008% of egress recordings processed in our US East region were terminated before completion. Affected recordings were marked EGRESS_FAILED ("egress timed out"). For file outputs, media was not written to the configured destination; for HLS and streaming destinations, media was delivered up to the point of failure. The cause was node scale-down evicting egress workers with active recordings. A fix was deployed at 21:00 UTC on 23 August.

Reported by LiveKit on their status page.

August 20, 2026

Investigating reports of elevated egress & connector API errors

Last update
  1. postmortemSep 10 · 19:08 UTC

    On August 20, 2026, customers in our US East region experienced approximately 10 minutes of 503 errors on the LiveKit Ingress API and failures launching media track egress, between 20:52 and 21:01 UTC. All other regions were unaffected.

    Root Cause
    An internal routing service lost one of its two instances, and a concurrent infrastructure issue with our cloud provider was preventing new compute nodes from provisioning in US East, so the replacement couldn't start. At 20:51 UTC, an automated maintenance process removed that last instance, making the service fully unavailable. It recovered at 21:01 UTC once the evicted instance freed capacity for a replacement.

    Timeline (UTC)

    • Aug 20, 20:52 - Automated maintenance removes last instance in the pool; customer-facing errors begin
    • Aug 20, 21:01 - Service restored
    • Aug 21, 03:00 - Cloud provider resolves the underlying infrastructure issue; region fully healthy

    Mitigations & Follow-ups

    • The service is restored with two instances guaranteed to run on separate physical nodes.
    • We updated the automated maintenance process to avoid evicting the last healthy instance of a service.
    • We added disruption protection for the affected service and are continuing broader resilience improvements.
    • Our cloud provider resolved the infrastructure issue that had prevented new nodes from scaling up.

Reported by LiveKit on their status page.

August 19, 2026

Investigating delayed session data ingestion on Cloud Dashboard

Last update
  1. resolvedAug 20 · 00:30 UTC

    This incident has been resolved. Between 19:20 and 00:10 UTC, session data and observability ingestion was delayed by 45-60 min. The root cause has been addressed and data ingestion returned to normal by 00:10 UTC. No data was lost during this window.

Reported by LiveKit on their status page.

August 17, 2026

Investigating reports of failed SIP Transfer calls in US East

Last update
  1. postmortemSep 1 · 11:44 UTC

    Summary

    On 17 August, between 17:24 and 18:05 UTC, a small percentage of SIP call transfers failed in our US East region during a routine configuration rollout. An issue in the automation that manages our SIP signaling servers prevented outgoing servers from being taken out of service safely, and those servers shut down while calls were still active on them. Transfers have completed normally since 18:05 UTC. Calls themselves stayed connected, and no other region was affected.

    Root Cause

    The configuration change restarts the servers that handle SIP signaling one at a time. Before a server is shut down, it is removed from service so that no new calls reach it, and it is then given some time to finish the calls it is already handling.

    In this case that removal did not complete expectedly and the outgoing servers kept receiving new calls for the entire drain duration, then shut down on schedule with calls still active on them.

    Timeline (UTC)

    • 16:52 - A routine configuration rollout begins in US East.
    • 17:24 - SIP transfer requests begin to fail for a subset of active calls.
    • 17:36 - Automated monitoring detects the elevated failure rate and our team begins investigating.
    • 17:40 - A further subset of active calls is affected.
    • 17:41 - Replacement capacity comes online.
    • 18:05 - Last of the impacted calls attempts a transfer and record a failure.

    Scope of Impact

    Only SIP call transfers (TransferSIPParticipant) in our US East region were affected, between 17:24 and 18:05 UTC. This represented 0.04% of all active calls in that window, and 1.1% of the calls that attempted a transfer. Customers who were impacted would have seen affected transfer requests return a 408; the underlying call stayed connected and only the transfer failed. Inbound and outbound calling were unaffected, as were calls that did not attempt a transfer, and no other region was affected.

    Mitigations and Follow-ups

    • We have deployed an alert for calls that end unexpectedly when a server shuts down.
    • We are changing our rollout process so that a server which cannot be removed from service safely halts the rollout.
    • We are preventing new calls from being routed to servers that are shutting down.
    • We are improving monitoring of the automation that manages SIP server rotation.
    • We are returning a more specific error when a transfer request cannot be delivered.

Reported by LiveKit on their status page.

August 14, 2026

Hosted agent builds failing in US East

Last update
  1. postmortemAug 17 · 17:14 UTC

    Summary

    An issue with the image registry which powers our hosted agents offering temporarily caused new hosted agent builds in US East to fail. Existing agent workloads were not affected, and no sessions or calls were impacted.

    Timeline (UTC)

    • 19:54 - Hosted agent builds begin failing in US East
    • 20:48 - Our team noticed the increased failure rate and began investigating
    • 21:04 - First mitigation deployed
    • 21:08 - Most builds succeeding
    • 22:14 - Second mitigation deployed, covering the remaining affected infrastructure
    • 22:24 - All builds succeeding

    Scope of Impact

    Only new hosted agent builds and deploys in US East were affected. Agents already running were unaffected, as were all sessions, calls, and other regions. Customers who were impacted would have seen their deploy fail with “unable to deploy agent: an error occurred while building your agent, please try again”.

    Mitigations and Follow-ups

    • We are improving alerting on build failures.
    • We are improving monitoring of our downstream image registries.

    We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

Reported by LiveKit on their status page.