Incidents | Thalassa Cloud Services Incidents reported on status page for Thalassa Cloud Services https://status.thalassa.cloud/ https://d1lppblt9t2x15.cloudfront.net/logos/a974135e3c164b63e0b01e926315d6c7.png Incidents | Thalassa Cloud Services https://status.thalassa.cloud/ en Kubernetes Node Machine Issues https://status.thalassa.cloud/incident/988311 Fri, 14 Aug 2026 09:06:00 -0000 https://status.thalassa.cloud/incident/988311#01d374dd2719bc1fe732938a25a273893402bf010ea583117bde2c1691925f52 On 31 July 2026, we deployed an updated Kubernetes node specification containing additional virtualization hardening measures. The specification change was unintentionally picked up by our automation as requiring node replacement, resulting in a platform-wide rollout. Safeguards were already in place to restrict these actions to scheduled cluster maintenance windows. Despite these safeguards, hundreds of Kubernetes nodes started being replaced simultaneously. Two issues occurred during the rollout that further delayed node replacements and caused temporary disruption for a small number of clusters. ## Root Cause ### Node Replacement We determine whether nodes require replacement by calculating a hash from relevant parts of the node pool specification. Safeguards and tests are in place to prevent configuration changes from unintentionally changing this hash. The new hardening setting defaults to `false` and was intentionally excluded from this calculation until enabled. The Kubernetes deployment layer normally enables the setting during the next scheduled maintenance window. However, deploying the updated specification also changed the instance type specification referenced by existing node pools. This caused the node pools to reconcile immediately and resulted in the new hardening configuration being applied outside the expected maintenance flow, triggering node replacements across the platform. ### Slow Node Provisioning Hundreds of nodes were provisioned and decommissioned within a short period. Some node drains were blocked by Pod Disruption Budgets (PDBs) or workloads taking longer to terminate. These drain operations occupied our node lifecycle controllers for up to 10 minutes per attempt, reducing the capacity available to provision replacement nodes. In some cases, new nodes took up to one hour to join a cluster. ## Remediation * Improved the node drain timeout and retry logic to prevent individual drain operations from blocking overall node lifecycle processing. * Added additional validation and safeguards to prevent unintended node replacements outside scheduled cluster maintenance windows. * Improved Kubernetes cluster and node testing in our staging environments to verify that nodes are not replaced outside their maintenance windows, in addition to the existing specification hash tests. Delays on volume provisioning, attach, detach actions https://status.thalassa.cloud/incident/1010811 Thu, 13 Aug 2026 13:23:00 -0000 https://status.thalassa.cloud/incident/1010811#959cdf2b388446a1864fddd1b6c378acd24f0af300151430d053e274a3b70afa All systems are operating at normal levels again. Delays on volume provisioning, attach, detach actions https://status.thalassa.cloud/incident/1010811 Thu, 13 Aug 2026 13:16:00 -0000 https://status.thalassa.cloud/incident/1010811#b5b4a092f1ac8c8c89ed95a8e35d02eae7f62b60ed972d332460b34d140aaa3d We are currently experiencing a large influx of operations on our volume provisioning system. We are investigating. Users may experience delays with volume provisioning, attaching, detaching or other volume operations. Availability of attached volumes is not impacted. Kubernetes Node Machine Issues https://status.thalassa.cloud/incident/988311 Fri, 31 Jul 2026 12:24:00 -0000 https://status.thalassa.cloud/incident/988311#2450c9b53f09476fbae953e10e0e7d5d0e38b71daefe96b991ac44769faa85d4 We will be publishing a Post Mortem once we complete our investigation and internal review, to share lessons learned and avoid this type of issue in the future. The post mortem will be shared on this page when available. Kubernetes Node Machine Issues https://status.thalassa.cloud/incident/988311 Fri, 31 Jul 2026 11:27:00 -0000 https://status.thalassa.cloud/incident/988311#1905e745fcf7dbfc88ed45711a4c47e9a367c584a970842b5bcf2a852d09705b This issue has been resolved. Kubernetes Node Machine Issues https://status.thalassa.cloud/incident/988311 Fri, 31 Jul 2026 11:11:00 -0000 https://status.thalassa.cloud/incident/988311#4c1bbccd6ab7923ce4ecfe903749dac2291867487bfe24aa8f054ccf00a3a285 Our monitoring indicates that the Kubernetes machines are all healthy. We are monitoring the situation to ensure we have not missed anything. Kubernetes Node Machine Issues https://status.thalassa.cloud/incident/988311 Fri, 31 Jul 2026 09:59:00 -0000 https://status.thalassa.cloud/incident/988311#7bec7d89110fe0833675e9ea2d2ebf876a397f3b67566f59bfde62df774a1743 A subset of clusters is experiencing delays with the provisioning of the replacement Kubernetes nodes. We are investigating. Kubernetes Node Machine Issues https://status.thalassa.cloud/incident/988311 Fri, 31 Jul 2026 09:49:00 -0000 https://status.thalassa.cloud/incident/988311#e3e4a2d2a4168641c9ef049c8f98d9d0567949513827c9ea02026776ddfaa7ac Due to security patches, a large amount of the Kubernetes Node Pools is receiving new Kubernetes machines. The security patch had an unexpectedly large impact, resulting in a large amount of Kubernetes Node Pools that are receiving new Kubernetes machines due to the virtual machine patches. Appologies for any inconvenience. Databases instances restart or switchover https://status.thalassa.cloud/incident/947021 Wed, 08 Jul 2026 09:28:00 -0000 https://status.thalassa.cloud/incident/947021#1c39c7e40f9f9d27b59123aaa190ea773799815e947c97d5cc1096e69c332453 We have identified the root cause and will be working on additional safeguards to protect against simliar issues in the future. Our systems indicate that all clusters are healthy. If you are still experiencing issues, please reach out to our support (https://helpdesk.thalassa.cloud/). Databases instances restart or switchover https://status.thalassa.cloud/incident/947021 Wed, 08 Jul 2026 09:22:00 -0000 https://status.thalassa.cloud/incident/947021#00ab96b2024a6859ff03c932684729760da9ad4606dac05899ce2b23805702ce Our systems indicate that the database instances are available again. This impacted customers running single instance clusters the most. If you were running more than 1 instance in your cluster, you experienced a switchover and had minimal impact. Databases instances restart or switchover https://status.thalassa.cloud/incident/947021 Wed, 08 Jul 2026 09:19:00 -0000 https://status.thalassa.cloud/incident/947021#7b35af5ce26f13d80061b5f104b93d4e595fee0621f379aadad346f8627e0708 We experienced an unexpected reboot of all database clusters while we were performing security patching. After the reboot, your database instance will be fully available again. Increased Timeouts errors on Object Storage uploads https://status.thalassa.cloud/incident/908681 Sat, 30 May 2026 07:45:00 -0000 https://status.thalassa.cloud/incident/908681#a278db6d650f04f7539519b967e1cad9813d08430230caca19b46fb42a8403a9 We have identified the issue and reverted a recent change that caused the increased timeouts during upload. Increased Timeouts errors on Object Storage uploads https://status.thalassa.cloud/incident/908681 Sat, 30 May 2026 07:40:00 -0000 https://status.thalassa.cloud/incident/908681#da9a21917db943c211d845525fb0c2dca47f86a35c1c472282dc57ec862fa034 We received reports that there was an increase in timeouts during uploads on our object storage. We are investigating. TFS Unreachable https://status.thalassa.cloud/incident/868688 Fri, 10 Apr 2026 12:06:00 -0000 https://status.thalassa.cloud/incident/868688#a788708093837bae9e8eb3e153ae706dd3661f16d92890375349bd51fa66d36e This incident has been resolved. TFS Unreachable https://status.thalassa.cloud/incident/868688 Fri, 10 Apr 2026 12:04:00 -0000 https://status.thalassa.cloud/incident/868688#86b28b6864efc5864d44da1a9fb1ed5da25856cb42ea397edf516b9ae31f9a30 We have identified the root cause and deployed a fix. We are seeing TFS instances becoming available again. TFS Unreachable https://status.thalassa.cloud/incident/868688 Fri, 10 Apr 2026 12:01:00 -0000 https://status.thalassa.cloud/incident/868688#8ff01eb52f22647453a0ee8e5c1639880f5c85f5a9741c8ce7bff45080e29d4b Some TFS instances may be unreachable at the moment. We are investigating. dgp instance types network issue https://status.thalassa.cloud/incident/859756 Sat, 28 Mar 2026 21:39:00 -0000 https://status.thalassa.cloud/incident/859756#84d7964c7bd319fa3bfc842b5cb9dc6e4a95aa06a1714f5f7b1e922deb494294 This issue has been resolved. An unintended network restart on some DGP hypervisors disrupted the block storage connection, causing several virtual machines to freeze. We initially migrated affected VMs, but they remained frozen. After manually unfreezing the instances, all services recovered. dgp instance types network issue https://status.thalassa.cloud/incident/859756 Sat, 28 Mar 2026 21:08:00 -0000 https://status.thalassa.cloud/incident/859756#8489293054bd0ae6fdd048322e4ac3cda8ac450ef7b2bcaf0508c712c9a8f5dd A subnet of the Virtual Machine Instances and Kubernetes Machine instnaces have been affected. We are working on restoring the services. Infrastructure Maintenance https://status.thalassa.cloud/incident/846014 Sat, 28 Mar 2026 21:00:00 -0000 https://status.thalassa.cloud/incident/846014#ebcae973eb041907a592b3cceee29f5794c120084febb20852a210c86ed839ce Maintenance completed dgp instance types network issue https://status.thalassa.cloud/incident/859756 Sat, 28 Mar 2026 20:44:00 -0000 https://status.thalassa.cloud/incident/859756#f3442eedef9dcebe41e681f06d9886c4e92bccd686355889b6415b5a4da85929 Some VMs and compute workloads running on dgp instance types have some network issues. We are working to investigate the issues. Infrastructure Maintenance https://status.thalassa.cloud/incident/846014 Sat, 28 Mar 2026 11:00:00 -0000 https://status.thalassa.cloud/incident/846014#d4fca4829e06acac0be5e8df139563cca0de806d2c4d310c780ef3148824be85 Thalassa Cloud will be performing scheduled extended maintenance for network, storage and hypervisor components on March 28th, starting at 12:00 CET. We aim to have no disruption to our services, however users may experience some increase in storage IO/Latency, network latency or reconnects on long-lived connections during this maintenance. NAT Gateways and Load Balancers https://status.thalassa.cloud/incident/829454 Fri, 20 Feb 2026 13:00:00 -0000 https://status.thalassa.cloud/incident/829454#031b668aa054912a2bea657b22f92bd522e611bd47b39b43aed955322b81abad We will be performing regular maintenance for our NAT Gateway and Load balancer services in order to improve our services. During this maintenance network, some connections may experience some packet drops as network components failover or are restarted.