Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

April 2025

Fixed · Global

We are investigating an issue with emailing of Clever Cloud Status.

Fixed · Infrastructure · Global

An hypervisor on the PAR region is unreachable and is currently rebooting. Services hosted on this hypervisor are also unreachable.

Fixed · Infrastructure · Global

An hypervisor on the PAR region was unresponsive between 11:48 CEST and 11:54 CEST. Applications on it are queued for redeployment. During this time, services on the hypervisor may have failed to respond or have elevated response time.

EDIT 12:53 CEST: All applications were redeployed before 12:12 CEST. The incident is now over.

March 2025

Fixed · MySQL shared cluster · Global

[16:12] (CET) After migrating all remaining addon on MySQL 8.0 MTL DEV cluster, we will shutdown it

Fixed · Access Logs · Global

We had an issue on the ingestion pipeline of access logs, we are recovering from the issue and consuming the lag. The lag will be fully consume by the end of the day.

Fixed · Global

We will migrate soon www.clevercloudstatus.com to a newer version of the underlaying software for a better experience !

No disruption should occur.

Please notify the support if you notice any issue.

Fixed · PostgreSQL shared cluster · Global

[15h26] (CET) After migrating all remaining addon on PostgreSQL 11 MTL DEV cluster, we will shutdown it

Issues on Grafana Alerting
Fixed · Metrics · Global

[13:30 CET] - Grafana instances failed after a maintenance restart

[14:50 CET] - Grafana is accessible, but alerting is deactivated: grafana alerting system failed restarting. We are currently investigating

[18:22 CET] - Grafana is accessible and alerting is reactivated: our proxy cause our database connection to failed during migration causing connection latency and multiple errors that prevent the Grafana to start.

[20:15 CET] - Grafana is accessible now in 9.5.13 - End of incident

Fixed · Pulsar · Global

09:47 UTC: After a routine upgrade of some components in a specific AZ, we've seen some disruptions in message ingestion 10

09:15 UTC: One of the component of broker was updated & rebooted

10:28 UTC: Message ingestion starts to recover and back to normal levels, incident solved

Fixed · Infrastructure · Global

Network Incident: Disrupted link with network provider

We would like to inform our users that a network incident was reported on one of our provider's links, which occurred last night. This incident has led to disruptions in access to some of our services.

Incident Details:

  • Type of Incident: Network connectivity issue=
  • Start Time: 01h30 UTC
  • Estimated End Time: 02h00 UTC

Impact:

Users may have experience increased response times or interruptions in access to certain services.

Fixed · Global

We are currently updating our controle plane and our images to support new MySQL versions 8.4.3-3 and 8.0.41-32

No impact expected

[16:20 CET] Those new version have correctly been released

Fixed · Global

We are currently updating our controle plane and our images to support new PostgreSQL image versions 16 and 17

No impact expected

[11:30 CET] This release has been correctly done

Fixed · Infrastructure · Global

An hypervisor is unreachable on the Paris region, we are investigating the issue

EDIT 21:15 UTC : A failed nvme is preventing the hypervisor to boot, but no customer services has been impacted. The hypervisor is only for testing purposes and distributed systems.

Fixed · Infrastructure · Global

An hypervisor is temporary overloaded on the MTL, we are investigating.

EDIT 12:00 UTC : the overload is caused by an hardware failure, we are draining the hypervisor. Databases on this hypervisor will migrate to another healthy hypervisor.

EDIT 12:45 UTC : we have migrated a few databases on the hypervisor, during the operation we had to reboot the hypervisor which resolved the issue, we are still investigating the root cause. Services on the hypervisor should be up.

Fixed · Infrastructure · Global

An hypervisor is rebooting on the MEA region, we are working to restart all services.

EDIT 16:44 CET: The hypervisor has rebooted and all services are available again.

Fixed · Infrastructure · Global

We are experiencing a network reachability issue on the Gouv region. We are looking into it.

EDIT 13:22 CET: Our infrastructure provider acknowledged the incident and is working on it.

EDIT 13:46 CET: Our infrastructure provider continues to investigate the issue.

EDIT 13:52 CET: The whole region is impacted, no service hosted on that can be reached.

EDIT 14:17 CET: A fix has been implemented on our infrastructure provider side. We regained access to the infrastructure since 14:08. We are making sure all services are restarted.

EDIT 14:32 CET: All services have been restarted, we keep watching if anything comes up.

EDIT 14:40 CET: All KMS nodes are fully operational

Fixed · Infrastructure · Global

Clever Cloud Incident – Explanations and Lessons Learned

Today we experienced an incident affecting our infrastructure. Here is a summary of the causes, ongoing analysis, and planned actions to strengthen our resilience.

Incident Timeline:

  1. Power Outage: A power failure reduced our computing capacity by one-third, highlighting the need to expand our infrastructure to five datacenters to better absorb such incidents.
  2. Network Issue: A network outage related to BGP announcements followed, revealing an underlying issue that requires further investigation to prevent recurrence.

Why Did Recovery Take Time?

  1. Machine Reconnection: The corrective measure to prevent overload during VM reconnection to Pulsar was not fully effective. An in-depth analysis is underway to improve this process.
  2. Orchestration Evolution: Our current system is reaching its limits. We are working on a new orchestration architecture to better manage recovery and optimize performance.

Next Steps:

We will publish a detailed post-mortem and schedule meetings with customers to:

  • Analyze the incident,
  • Explain upcoming changes,
  • Demonstrate our commitment to improving infrastructure resilience.

We will keep you informed about these actions. Thank you for your patience and trust.

The Clever Cloud Team

Detailed timeline :

EDIT 2:11pm (UTC): One of the Paris datacenters has experienced an electricity issue. Some hypervisors have been rebooted.

EDIT 2:21pm (UTC): Applications are being restarted, the situation is stabilizing.

EDIT 2:28pm (UTC): Monitoring is up and generating new statuses

EDIT 2:34pm (UTC): Recovering process is still ongoing. Customers can open a ticket though their email endpoint in addition to the Web Console.

EDIT 2:43pm (UTC): Infrastructure is under high load. We're accelerating the recovery process with load sanitization

EDIT 2:48pm (UTC): All Load balancers are now back in sync. Per service Availability :

  • Load Balancing (dedicated instances) : OK
  • Cellar Storage : OK
  • Orchestration : OK but recovering a lot of runtime instances.
  • API : OK
  • Metrics API : KO (network topology split)
  • Databases : Globally OK (individual situations being worked on)
  • Monitoring : OK (in sync from a couple of minutes)
  • Infrastructure overall load : high

EDIT 3:00pm (UTC): Deployments are done but still slow. Clever Cloud API is being restarted.

EDIT 3:05pm (UTC): Infrastructure load sanitization : 30 %

EDIT 3:40pm (UTC): Overall situation is better. Still a few tousands VMs in the recovery queue.

EDIT 3:48pm (UTC): Most databases should be available. We're experiencing an additional delay with some encrypted databases.

EDIT 3:53pm (UTC): Infrastructure load sanitization : 100 %

EDIT 4:15pm (UTC): Remaining Databases are recovered. Estimated total fix time is expected in the 10/15 next minutes.

EDIT 4:30pm (UTC): All applications should run fine (orchestration point of view). Orchestration monitoring is partially up and running (stuck apps will be quickly unstuck)

EDIT 5:10pm (UTC): All stuck applications should now be available (monitoring point of view).

EDIT 5:16pm (UTC): All databases are available

EDIT 5:44pm (UTC): Incident is considered closed (a few cases are still dealt with customers)

February 2025

Fixed · Infrastructure · Global

An Hypervisor in the PAR region is unreachable, we are working on it.

EDIT 16:30 UTC+1: This hypervisor has faced a network issue making it unreachable during several minutes. It is now up and running.

EDIT 19:15 UTC+1: The incident is over.

Fixed · MySQL shared cluster · Global

[2025-02-27] 12:16UTC We are deploying a new shared cluster for Dev mysql 8.4 addon

[2025-02-27] 11:40UTC New mysql cluster successfully deployed

[2025-02-27] 13:30UTC We are deploying a new shared cluster for Dev postgres 15 addon

[2025-02-28] 09:17UTC New postgresql cluster successfully deployed

Fixed · Reverse Proxies · Global

We made a renewal of wildcard certificate *.services.clever-cloud.com and some services were impacted. We are investigating the issue and had roll back the new certificate.