
Grafana Cloud Incident History
Grafana Cloud is currently operational with all systems functioning normally.
Incident History
Showing incidents from the last 15 days
Report: "Write Outage"
Last updateThis incident has been resolved.
Healthy as of 17:00 UTC. We are continuing to monitor.
We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.
Report: "Errors creating new Slack integration for Grafana IRM"
Last updateThis incident has been resolved.
We have verified the fix, and we are starting to roll it out
We have identified the root cause of the problem, and we are working on a fix
We are investigating a possible error creating Slack integrations. Will update the status soon.
Report: "Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23"
Last updateWrites are looking stable in the last 6h
status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor
Report: "Metrics Write Path Errors"
Last updateThis incident has been resolved.
Error rates have dropped, and we are monitoring this issue for any recurrence.
We are currently investigating an elevated rate of error and latency in the impacted write path.
Report: "Stacks using SCIM user provisioning are currently unable to log into Grafana"
Last updateThis incident has been resolved.
We have applied a fix and are currently awaiting feedback from affected customers.
The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
We are currently facing an issue where Stacks using SCIM user provisioning are currently unable to log into Grafana. We are currently investigating this issue and working on a fix.
Report: "K6 - Cloud output test-runs are failing to fetch the script logs"
Last updateThis incident has been resolved.
We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.
Report: "Cloud Log Exporter Unavailable"
Last updateThis incident has been resolved.
A fix has been implemented and we are monitoring the results.
Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.
Report: "PDC Issues"
Last updateThis incident has been resolved.
A fix has been implemented, and we are observing recovery. We will continue to monitor the results.
We are continuing to investigate this issue.
We are currently investigating an issue that is causing issues with PDC in the prod-eu-west-2 region. We will provide another update in 1-2 hours.
Report: "Issue with Dashboard Views Being Registered"
Last updateA fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.
We are currently working on a fix related to this incident. There are no new updates to share at this time.
We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
We've identified the cause of the issue impacting dashboard view and error counts in the Dashboards and Folder list. Our team is currently working on implementing a fix and validating the solution. We will provide another update within the next 24 hours, or sooner if we have additional information to share.
This is related to https://status.grafana.com/incidents/rhrk2ck6ly0y which was resolved by mistake. For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly. This is ultimately causing inaccurate or missing view counts. We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
Report: "Grafana Cloud login issues for users with role of None"
Last updateUsers with the None role authenticating via Grafana.com should be able to access their stacks again
We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
We’re currently investigating an issue with user logins with the None role in Grafana Cloud. Our team is actively working to identify the cause. Thank you for your patience.
Report: "Adaptive Metrics aggregation delay in eu-west-0 region."
Last updateThe incident is now fully resolved.
Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.
We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the appropriate fix is already deployed. We're seeing services recovering. More updated to come soon.
Report: "Delayed Aggregated Metrics (prod-us-central-0)"
Last updateThis incident has been resolved.
We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.
We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively working on a mitigation. We will provide further updates as the investigation progresses.
Report: "Partial Outage in prod-eu-west-2"
Last updateThis incident has been resolved. Thank you for your patience.
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
We are continuing to investigate this issue.
We are investigating a broader issue affecting multiple Grafana Cloud products in prod-eu-west-2. This appears to be caused by a third-party provider issue rather than a Loki-specific problem. Impact is currently inconsistent: some components are affected while others continue to function normally. We have also seen some impact to Tempo write paths. We are continuing to investigate and will share another update as soon as we have more information.
We are investigating an issue affecting Loki queries in prod-eu-west-2. We first observed this behavior at approximately 21:57 UTC. Affected users may see elevated query errors, timeouts, or intermittent failures when running Loki queries in this region. The issue appears to be improving, but it is not fully resolved yet. We are continuing to investigate and will share another update as soon as we have more information.
Report: "Some Reports of Grafana Not Loading."
Last updateThis incident has been resolved. Thank you for your patience.
The fix is in the process of being rolled out.
We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
We are currently investigating an issue impacting a small subset of stacks in the impacted regions. We will provide more details as they become available.
Report: "Fleet Managment Interface 404's"
Last updateBetween 9:45 and 12:00 UTC, we experienced an issue affecting the Fleet Management interface. During this time, the Remote Configuration tab for production stacks returned a 404 error when accessed through the UI. Backend systems were not affected. Configuration-as-code workflows and other non-UI methods of interacting with Fleet Management continued to operate normally throughout the incident. This issue has since been resolved.
Report: "Mimir Write Performance Degradation"
Last updateThis incident has been resolved. Thank you for your patience.
From 19:24 until 19:40 UTC, backend performance degradation impacted ingestion. We are currently monitoring.
Report: "Mimir Partial Write Outage"
Last updateThis incident has been resolved.
The outage is recovered as of 16:56 UTC, and we're continuing to monitor.
We're investigating backend degradation which has resulted in a partial write outage beginning around 16:40 UTC. This degradation has also impacted reads and rule evaluation. Issue has been identified and we are working on mitigation.
Report: "IRM Mobile App forcing some users to logout"
Last updateThis incident has been resolved.
A new iOS mobile app version (v2.39.5) has been released, which includes a fix for this issue. We are monitoring the rollout and its impact to ensure the issue has been fully resolved. If you continue to experience any problems after updating to the latest version, please let us know.
A fix has been submitted for review and will be released as Grafana Mobile v2.39.5 once it is approved. If your app is currently on v2.39.3, we recommend not upgrading to v2.39.4 and instead waiting for v2.39.5 to become available. Users who have already been signed out by this issue can log back into the app and continue using it. We will provide another update once v2.39.5 is available.
Our investigation has narrowed the impact to version 2.39.4 of the Grafana mobile app on iOS. Users should avoid upgrading to this version until an updated release is available. As a precaution, we also recommend ensuring you have alternative notification methods configured (such as SMS, email, or phone calls) if you rely on mobile push notifications for alerting. We are actively working on a resolution and will provide another update as soon as more information is available.
We are aware of an issue in the latest mobile app release that is causing some users to be signed out and asked to log in again. We have reproduced the behavior and are investigating and working on a fix. We will share another update as soon as we have more information.
Report: "Issue with Dashboard Views Being Registered"
Last updateAt this stage, we are considering the incident resolved.
We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly This is ultimately causing inaccurate or missing view counts.