Service Disruption for a subset of customers ANZ
Last updatepostmortemSep 15 · 01:15 UTC
Incident Summary
Between 31 August 2026 and 2 September 2026, some customers experienced intermittent service disruptions, including slow application responsiveness, connection timeouts, and periods of Ci unavailability. The issue was caused by internal DNS resolution failures within an underlying CiA cluster, which affected communication between application services and backend caching components. Service stability has been restored and all affected environments are operating normally.
Root Cause
A significant increase in internal DNS requests related to Redis caching placed unexpected demand on the cluster's DNS service (CoreDNS). This resulted in memory exhaustion within CoreDNS, preventing it from scaling effectively to meet demand. As a result, some application services experienced delayed or failed DNS lookups, leading to connection timeouts and degraded application performance.
Corrective Actions Taken
To restore service and stabilise the platform, we:
- Increased CoreDNS capacity by scaling the service.
- Rebalanced workloads by moving affected environments to alternative CiA clusters.
- Restarted impacted DNS services and closely monitored DNS performance until normal operation was confirmed.
Preventative Actions
- Optimising Redis connection management and DNS resolution behaviour to reduce DNS query volumes.
- Reviewing and refining CoreDNS autoscaling policies, resource limits, and capacity thresholds.
- Assessing long-term cluster capacity and workload distribution strategies.
- Enhancing monitoring and alerting to identify and respond to DNS-related issues before customer impact occurs.
Reported by TechnologyOne on their status page.
