Shops not available
Last updatepostmortemAug 27 · 14:38 UTC
## Summary On 21 August 2026, a storage backend incident at our datacenter provider caused an outage affecting slightly more than half of the shops hosted on our platform. Our monitoring detected the first service problems at **12:59:57 EEST**, and investigation started immediately. We worked closely with our datacenter provider throughout the incident while they investigated the storage backend, verified the affected systems and performed the manual actions required before the servers could be safely started again. All affected Vilkas application servers were back online by approximately **16:33 EEST**. We continued enhanced monitoring afterwards to ensure the environment remained stable. ## Root cause According to our datacenter provider, the root cause was a previously undetected hardware defect affecting a specific storage hardware model. The storage platform is redundant, with server data stored on multiple physical storage systems. Under normal circumstances, an issue with one storage backend is handled automatically by continuing operation through another backend. In this incident, however, some virtual servers entered an error state and required manual verification before they could be safely started again. This meant that resolving the underlying storage issue alone was not sufficient. The state of the affected storage and servers also had to be verified before services could be brought back online in a controlled manner. ## Timeline * **12:59:57 EEST** – Our monitoring detected the first service problems. Investigation started immediately. * **13:05** – Our datacenter provider identified an issue in its storage environment. * **13:12** – The provider confirmed the affected infrastructure and escalated the incident internally. * **13:23** – The provider had started work on the affected storage systems and confirmed to us that our infrastructure was impacted by the storage incident. * **13:27** – We published this incident on our status page. * **13:38** – The provider identified virtual servers that had entered an error state as a result of the storage incident. * **13:59** – Manual verification and actions on affected servers were in progress. * **During the afternoon** – We remained in continuous contact with the datacenter provider while the storage state and affected servers were verified and we coordinated the safe startup of our infrastructure. * **16:20** – The affected storage server was available again. * **16:33** – All affected Vilkas application servers were running. * **16:49** – Normal service operation had been confirmed and enhanced monitoring continued. ## Data integrity We have found **no indication of data loss or corruption** related to the incident. Affected shops were intentionally kept unavailable while the storage environment and affected servers were being verified. Services were only brought back online once the required systems could be started safely and in a controlled manner. ## Preventive actions Our datacenter provider has identified and mitigated the hardware defect responsible for the incident. They have also stated that they are improving their storage failover software as well as their internal processes and tooling for handling similar storage incidents. We are also reviewing the incident internally, including our startup procedures, monitoring and incident communication. We apologize for the disruption and thank you for your patience while we worked with our datacenter provider to bring the affected services safely back online.
Reported by Vilkas on their status page.
