Aptible logo and current status indicator

Aptible Incident History

Aptible is currently operational with all systems functioning normally.

Last checked Jul 28, 2026 5:38 AM UTC from Aptible's official status page

Incident History

Showing incidents from the last 15 days

Report: "Operation failures due to AWS Outage"

Last update
resolved

This incident has been resolved.

monitoring

We have applied a mitigation and re-enabled operations in us-east-1. We are monitoring for stability and will resolve this incident once we confirm conditions remain stable.

investigating

Operations remain disabled in us-east-1 while we continue to investigate AWS networking issues in the region. We are testing mitigations and will restore operations as soon as it is safe to do so. We will provide another update no later than 10:00 AM Eastern Time.

investigating

AWS is experiencing an issue with their APIs in us-east-1, and many services are impacted at this time. This may lead to failed operations. To avoid further issues, we have disabled operations in this region while we understand the full impact.

Report: "Delayed and Timed-out requests to the Aptible LLM Gateway"

Last update
resolved

Following the fixes deployed per the previous update, we have seen stable performance on the LLM Gateway since 7:12 EDT and are resolving the issue. The root cause of this issue was a severe performance degradation in 3 specific endpoints: • `GET /models` • `GET /v1/models` • `GET /metadata/models` ...all of which return a list of available models. The performance issue on these endpoints was severe enough to cause significant request queueing and degraded performance for other unaffected endpoints. The fix we deployed added caching on this endpoint to ensure fast responses. Since those fixes have been deployed, we are no longer seeing request queueing or impacts to other LLM Gateway API endpoints.

monitoring

After a regression in performance starting at 9:57 UTC (5:57 EDT) and lasting until 11:12 UTC (7:12 EDT), we've deployed a potential fix and updated our service scaling. We are continuing to monitor whether these changes fully resolve the issue before closing this incident.

monitoring

We have increased resource capacity, and we are monitoring the results.

identified

We are aware of delayed and timed-out requests to the LLM Gateway. We have made some changes to help minimize the impact while engineers continue to address the root cause.