
Coveralls Incident History
Coveralls is currently operational with all systems functioning normally.
Incident History
Showing incidents from the last 15 days
Report: "Increased latency for large repos"
Last updateThis issue was resolved over the weekend. We will continue to monitor for elevated latency in the modified queues.
Latency remains elevated by about 30% for large repos. We are continuing to monitor and do what we can to mitigate until we can implement a longer-term solution.
We are continuing to proactively monitor and clear the large repo queues. We are working on a longer term solution over the weekend and will post about it in the postmortem for this issue when complete. We will keep this incident open as long as we have large repo queues with above-average latency.
We are continuing to monitor for any further issues.
We have reduced traffic in affected queues by 75%. We still expect to clear the queue within about 1 hr. We will post updates here. If you think you may have been paused as a user with a large repo (>=5K source files) that has uploaded 3-10x normal daily uploads, please reach out to us at support@coveralls.io and we'll confirm and give you some next steps to avoid further pauses (see Usage Add-Ons at https://coveralls.io/pricing).
All systems fully functional. We have identified increased latency for large repos again and have proactively arrested sources of excess requests. We expect to restore normal latency within the hour.
Report: "Increased latency for large repos"
Last updateThis incident has been resolved but we believe it triggered a worsened incident overnight Sun night/Mon morning (US PDT) which has just been resolved. To be confirmed by full RCA, we believe a deluge of large repo uploads tied up individual web servers that handle frontline requests. Each web server is able to recover on its own, and did, but as volume increased all servers were eventually affected, only allowing short windows where requests could get through—rejecting most requests with 504 errors. We'll post a post-mortem when we understand more about what happened and how to prevent it going forward.
We are continuing to monitor this situation. Latency is much reduced but still elevated. We will close this incident when it's fully restored to normal.
We are monitoring increased latency for larger repos (5K+ source files) due to elevated traffic in background job queues. We will pause outlier repos (more than 2x average traffic) to allow queues to clear for the general population, then restore once queues have cleared. If you think your repo may have been one of those paused, two things: 1) Reach out to us at support@coveralls.io and we'll confirm; and 2) Consider purchasing a usage add-on so that your jobs are not throttled for fair-use, or an enterprise plan with isolated infrastructure for maximum performance: https://coveralls.io/pricing