Back to overview
Resolved

Kubernetes Node Machine Issues

Jul 31, 2026 at 9:49am UTC
Affected services
Kubernetes Service

Resolved
Aug 14, 2026 at 9:06am UTC

On 31 July 2026, we deployed an updated Kubernetes node specification containing additional virtualization hardening measures. The specification change was unintentionally picked up by our automation as requiring node replacement, resulting in a platform-wide rollout.

Safeguards were already in place to restrict these actions to scheduled cluster maintenance windows. Despite these safeguards, hundreds of Kubernetes nodes started being replaced simultaneously.

Two issues occurred during the rollout that further delayed node replacements and caused temporary disruption for a small number of clusters.

Root Cause

Node Replacement

We determine whether nodes require replacement by calculating a hash from relevant parts of the node pool specification. Safeguards and tests are in place to prevent configuration changes from unintentionally changing this hash.

The new hardening setting defaults to false and was intentionally excluded from this calculation until enabled. The Kubernetes deployment layer normally enables the setting during the next scheduled maintenance window.

However, deploying the updated specification also changed the instance type specification referenced by existing node pools. This caused the node pools to reconcile immediately and resulted in the new hardening configuration being applied outside the expected maintenance flow, triggering node replacements across the platform.

Slow Node Provisioning

Hundreds of nodes were provisioned and decommissioned within a short period. Some node drains were blocked by Pod Disruption Budgets (PDBs) or workloads taking longer to terminate.

These drain operations occupied our node lifecycle controllers for up to 10 minutes per attempt, reducing the capacity available to provision replacement nodes. In some cases, new nodes took up to one hour to join a cluster.

Remediation

  • Improved the node drain timeout and retry logic to prevent individual drain operations from blocking overall node lifecycle processing.
  • Added additional validation and safeguards to prevent unintended node replacements outside scheduled cluster maintenance windows.
  • Improved Kubernetes cluster and node testing in our staging environments to verify that nodes are not replaced outside their maintenance windows, in addition to the existing specification hash tests.

Updated
Jul 31, 2026 at 12:24pm UTC

We will be publishing a Post Mortem once we complete our investigation and internal review, to share lessons learned and avoid this type of issue in the future.

The post mortem will be shared on this page when available.

Updated
Jul 31, 2026 at 11:27am UTC

This issue has been resolved.

Updated
Jul 31, 2026 at 11:11am UTC

Our monitoring indicates that the Kubernetes machines are all healthy. We are monitoring the situation to ensure we have not missed anything.

Updated
Jul 31, 2026 at 9:59am UTC

A subset of clusters is experiencing delays with the provisioning of the replacement Kubernetes nodes. We are investigating.

Created
Jul 31, 2026 at 9:49am UTC

Due to security patches, a large amount of the Kubernetes Node Pools is receiving new Kubernetes machines.

The security patch had an unexpectedly large impact, resulting in a large amount of Kubernetes Node Pools that are receiving new Kubernetes machines due to the virtual machine patches. Appologies for any inconvenience.