Jobvite Access Issues
Last updatepostmortemAug 24 · 19:04 UTC
Incident Summary
between approximately 8:06 AM and 9:04 AM EDT, some users may have experienced difficulty logging in or intermittent unavailability while our team completed a routine production deployment.
Detection
The issue was identified internally during deployment verification and escalated immediately to internal teams.
Root Cause
A server instance involved in the deployment was unexpectedly replaced by our cloud provider. While a replacement was provisioned automatically, restoring it to active service required a manual step, extending the disruption.
Resolution
Engineering manually restored the replacement server to active service, returning the system to normal operation by 9:04 AM EDT.
Preventative Measures
- Automation: Automating the step that currently requires manual intervention to restore a replaced server to service.
- Monitoring: Adding alerting for server replacement events and for disk/CPU saturation on affected servers, and reconciling deployment-pipeline status against actual deployment outcome so failures cannot go undetected.
- Housekeeping: Implementing log rotation on affected servers to prevent disk space exhaustion.
Reported by Jobvite Operational on their status page.
