Monitoring - Fixes to scale worker capacity and improve service autoscaling have been deployed. We are observing scheduled calculations resuming execution. Due to the backlog accumulated during the degradation, jobs are currently catching up and may experience processing delays over the next few hours. We are actively monitoring the service until all scheduled calculations are fully up to date.
Monitoring
Monitoring - Fixes to scale worker capacity and improve service autoscaling have been deployed. We are observing scheduled calculations resuming execution. Due to the backlog accumulated during the degradation, jobs are currently catching up and may experience processing delays over the next few hours. We are actively monitoring the service until all scheduled calculations are fully up to date.
Investigating
Current Status Update — INC-3602 (Scheduled Calculations in Charts)
Impact:
Scheduled calculation executions in Charts are failing or significantly delayed, preventing new data points from being written. The issue currently appears isolated to cluster az-tyo-gp-001 (verified working normally on aw-tyo-001).
Current Findings:
High CPU utilization (100–1000%) observed on backend workers, causing jobs to fall hours behind schedule.
Extensive timeout errors and ~236 failed jobs recorded across the cluster over the weekend.
Newly created and existing scheduled calculations are affected.
Current Actions & Next Steps:
Engineering is investigating the root cause of the CPU spike and timeout failures.
Evaluating remediation options to relieve backend load and process the backlog of calculations.
Investigating
We are currently investigating an issue where scheduled calculations created in Charts are delayed or failing to write new data points. Our engineering team is actively investigating the root cause. Further updates will be provided as soon as they become available.