Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

November 2020

Fixed · Infrastructure · Global

We have detected a network incident between Free (French ISP) and our network provider for the Paris zone. We are seeing 40 to 50% of packet loss on this interconnection.

Our network provider is investigating the issue.

10:48 UTC: We no longer experience packet loss on this interconnection. We are awaiting more information from our network provider on the cause and resolution of this incident.

10:57 UTC: The issue is back, we are experiencing the same amount of loss again.

11:07 UTC: The issue went away again. We are still awaiting word from our provider.

11:32 UTC: We are experiencing packet loss again on the same link.

11:35 UTC: The issue went away again.

11:36 UTC: The issue, ultimately, lies with Free and we cannot do anything about it from our side. Until the root cause is properly fixed, the loss issue may come back off and on.

14:53 UTC: Our network provider tells us that the peering link has been affected by the side effects of a DDOS targeting another customer of our network provider. They are working on providing measures to prevent more attacks targeting this network which should in turn prevent this link from getting overwhelmed.

October 2020

Deployments timing out
Fixed · Deployments · Global

We are seeing some deployments timing out. It looks like the retry mechanism is doing its job just fine and deployments are starting anyway for all affected applications but you may be observing an usual delay.

We are investigating this issue.

13:47 UTC: The issue is fixed. All deployments have been working fine during this period, only delayed by a few seconds. The issue came from a misconfigured deployment component which was sending broken messages to hypervisors. The broken component has been dealt with.

Small Network downtime
Fixed · Infrastructure · Global
  • A core RabbitMQ node stopped responded and some databases were unreachable for 30 seconds. We are investigating the outage.
  • Some applications may register a connection loss to their database.

13:17 UTC - no other network loss. All critical parts of Clever Cloud have been checked and restarted to make sure they still communicate with each other.

Logs interuption
Fixed · Services Logs · Global

Logs were interrupted for 15 minutes due to an internal issue. They were recorded and are being ingested. It may take a few minutes to receive all logs and current logs.

The issue is currently fixed and awaiting for full resolution.

Fixed · FS Buckets · Global

Some FS Buckets addons are experiencing issues. We identified the issue and are working on its resolution.

EDIT 21:25 UTC: The issue is fixed. The PHP applications may not work correctly. We are redeploying them.

EDIT 22:30 UTC : Applications with FS Buckets have been redeployed. The incident is closed.

Post mortem: An incorrect human action conducted the FS Buckets system to follow the wrong path between different storage nodes. We applied fix to avoid this cause.

Fixed · Cellar · Global

We identified issues on Cellar addons availability. We identified them and are working on their resolution.

EDIT 15:25 UTC: fixed. We are investigating the reasons.

EDIT 15:45 UTC: we identified the reasons and applied a fix.

Fixed · Access Logs · Global

Metrics and access logs requests might experience issues following the maintenance of a core component of those features. Requests can either take a very long time to complete or simply answer an error. We are working toward a fix.

Data won't be lost, the ingestion is simply delayed.

Impacted products:

  • Metrics (in console or using the API)
  • Access logs (charts in the console's overview or using the CLI / API)

EDIT 14:03 UTC: Ingestion is now catching up on the delay, everything looks good. Looks like it may take 30 to 40 minutes to go completely back to normal.

EDIT 14:25 UTC: Ingestion has now caught up, everything should be back to normal.

EDIT 21:26 UTC: New issues are ongoing, we are investigating.

EDIT 22:16 UTC: Ingestion is running. We are consuming queues.

EDIT 23:30 UTC: Ingestion is back to normal. Fixed.

Fixed · Console · Global

Loading the console might result in various errors preventing users from logging in. We are currently investigating. The CLI shouldn't be impacted. Already loaded console webpages shouldn't be impacted either.

EDIT 13:02 UTC: A change causing this issue has been backed out. We will investigate further why it went wrong despite working correctly on our test infrastructure. Sorry for the disruption.

Fixed · Redis · Global

The redsmin dashboard for redis add-ons is currently unavailable. The Redsmin provider has been notified. We will update this post as soon as we have an update.

EDIT 10:54 UTC: Redsmin is currently working on a fix.

EDIT 19:54 UTC: The fix seems to be complete. Redsmin interfaces should now be able to load.

September 2020

Fixed · Services Logs · Global

There was an issue with regards to new logs collection between 16:45 and 17:00 UTC Some of these logs may have taken more time than usual to be processed. No logs have been lost.

Fixed · PostgreSQL shared cluster · Global

0727 UTC: The free shared postgresql cluster Leader has crashed due to disk issue 1000 UTC: The team sees the issue. 1016 UTC: The team promotes the follower as leader. 1050 UTC: All applications using dbs on that cluster are redeployed.

Login issues
Fixed · API · Global

We are currently looking into a login issue. Once you validated the form, the login process will reset, not allowing you to proceed to the wanted resource (console / CLI / other).

For any support queries, you can send us an email at support@clever-cloud.com

EDIT 14:26 UTC: The issue has been found and should now be fixed. We will investigate it further to prevent it from happening again.

August 2020

Fixed · Cellar · Global

Our old Cellar cluster (cellar.services.clever-cloud.com) which still has some data nodes on Scaleway is currently unreachable due to networking issues on Scaleway's side: https://status.scaleway.com/incident/956

We are monitoring the situation. Our new Cellar cluster (cellar-c2.services.clever-cloud.com) is still reachable and works fine.

EDIT 12:02 UTC: A reverse proxy node is somehow still able to communicate with the nodes on Scaleway. All cellar-c1 traffic has been routed through that reverse proxy and requests should be served as expected.

EDIT 12:34 UTC: The network issue seems to not be on Scaleway's side per say but more on Level3/CenturyLink side which is a more global networking provider.

EDIT 15:17 UTC: The incident on Level3/CenturyLink seems to be resolved. The cluster is now fully reachable.

Fixed · PostgreSQL shared cluster · Global

Postgresql-c1 which is an old PostgreSQL cluster still hosted on Scaleway may currently be unreachable due to some Level3/CenturyLink networking issues. Scaleway has an incident opened here: https://status.scaleway.com/incident/956

EDIT 15:17 UTC: The incident on Level3/CenturyLink seems to be resolved. The cluster is now fully reachable.

Fixed · Infrastructure · Global

Due to an outage of the Level3/CenturyLink networking provider, you might experience issues:

  • reaching our services: if your FAI uses this provider, you might experience timeouts reaching our infrastructure

  • reaching external services from our infrastructure: if you contact external services from our infrastructure, the peering routes might use this network provider and your requests might timeout too.

This incident will group the previous opened incidents:

  • https://www.clevercloudstatus.com/incident/294

  • https://www.clevercloudstatus.com/incident/295

We do not have an ETA for the service to come back to normal.

EDIT 15:17 UTC: The incident on Level3/CenturyLink seems to be resolved. All connections either incoming or outgoing to/from our services should be working as expected. Please reach to our support if not.

Fixed · Infrastructure · Global

An hypervisor went down (electrically shut off) unexpectedly.

This was caused by a human error, partly related to a laggy UI (low-level UI of a server manager used for a group of servers).

The person who triggered this realized the issue immediately and restarted the server which has stopped responding to our monitoring for a total of 3 minutes.

Chronology:

14:01:30 UTC: The server goes down

14:04:30 UTC: The server responds to our monitoring again and starts restarting static VMs (add-ons and custom services)

14:07:05 UTC: The last static VM starts answering to our monitoring again.

Impact:

Customers with add-ons on this server will find connection errors in their application logs during those 3 to 6 minutes and those applications most likely responded with errors to end users during that time.

Customers with applications with a single instance which happened to be on that server will have experienced about 2 to 3 minutes of downtime before a new instance started responding on another server.

Access logs unavailability
Fixed · Global

New access logs are currently not processed. They are currently kept until they can be processed again. Access logs emitted before this incident are fine.

This impacts:

  • Access logs fetch using the CLI or the API
  • Live request map in the console
  • Total number of requests / status codes in the console (those are still available to display but the total of requests will be wrong in a few hours as access logs emitted won't be taken into account)

The issue has been identified and we are working toward a fix.

EDIT 14:07 UTC: The problem has been solved and the access logs stored have been processed. You should now be able to have an up-to-date livemap and fetch recent access logs using the CLI / API. Request count will be affected and won't be computed for the time window the access logs were not processed.

Fixed · Access Logs · Global

The Metrics platform is unavailable at the moment. We are investigating the source of the issue.

14:30 UTC: It looks like an issue with the storage backend, we are working on bringing it back to life.

14:52 UTC: The storage backend looks fine but writes are still failing. We are still investigating this issue. It may take a while.

15:11 UTC: Again, the storage backend looked perfectly fine... restarting everything did fix the issue though so then again maybe it wasn't fine after all. Writes are functional, ingestion is working at full speed, fresh data will be available in ~20 minutes.

15:30 UTC: Ingestion delay is back to normal. Incident is over.

Fixed · SSH Gateway · Global

SSH connections may have failed randomly those past few hours. The root cause has been found and a better monitoring will be put in place. Instances should have automatically reconnected after a restart of one of the main components. If it didn't, you can try to restart your application. If you absolutely need to SSH to your application to debug something before restarting it, ping us on the support

July 2020

Fixed · Access Logs · Global

The Metrics platform is unavailable at the moment due to an issue with the storage backend. We are investigating.

16:23 UTC: Some storage nodes were misbehaving. The issue is now fixed: reads are functional again and ingestion is now catching up.

16:28 UTC: Ingestion delay has been divided by two, incident should be over in under 10 minutes.

16:28 UTC: Incident is over.