Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

October 2019

Fixed · Access Logs · Global

We are experiencing an issue on the Metrics service which is due to an error while adding capacity to the storage cluster. We are working on it.

10:26 UTC: The ingestion issue is fixed, the system is now catching up.

10:33 UTC: The ingestion delay is almost back to normal.

10:36 UTC: There is still a bit of a lag but it should come back to normal in a few minutes. Read performance is still a bit hit or miss but coming back to normal as well. We will reopen the incident if it does not.

11:06 UTC: The ingestion lag is increasing. We are investigating. This may take a while.

11:30 UTC: The cause has been identified and partially fixed.

11:37 UTC: Lag is now <5s ; we are currently working on fixing the issue in a more permanent way.

11:45 UTC: The issue is now fixed.

Logs collection issue
Fixed · Services Logs · Global

We have an issue with logs collection in the Paris zone. We are working on it.

13:20 UTC: The issue has been identified and at least partially fixed. Logs are coming through but we are still making sure that everything is indeed fine.

13:25 UTC: The issue is indeed fixed. Some older logs are still being collected.

13:33 UTC: Incident is over.

September 2019

Fixed · MongoDB shared cluster · Global

A free shared mongodb cluster has too many connections opened which prevents new connections from working. We are looking into which user(s) are opening too many connections and we will start a new cluster to alleviate the issue. We have no immediate solution, sorry for the inconvenience.

08:40: The problem has been alleviated by allowing more connections. It will slow down the service but you can at least connect to your databases and migrate to paid add-ons if you were using this service for production. We will start a new cluster very soon to improve performance.

Fixed · Deployments · Global

False positives in monitoring are causing a lot of deployments, making legit deployments harder to process.

17:21 UTC: Incident is over. A monitoring component was still complaining about a few applications in a loop, there was no actual issue, just a very overzealous alerter process. Deployments performance has been back to normal since 16:43 however.

Deployments delayed
Fixed · Deployments · Global

Deployments are delayed because of an unusual amount of deployments to be processed.

12:44 UTC: The delay is now back to normal. Some deployments may be stuck though, please contact us if you are experiencing such an issue.

Fixed · Global

There seems to be a problem with download and upload of build cache archive causing them to hang. It is resolved for now but we are watching closely to see if the problem reappears

Logs experiencing issues
Fixed · Services Logs · Global

Our logging infrastructure (including live logs) is experiencing issues.

EDIT 20:51 UTC: fixed.

EDIT 23:19 UTC: the logging infrastructure is experiencing issues. We are working on a fix.

EDIT 23:25 UTC: fixed.

August 2019

Fixed · MySQL shared cluster · Global

The cluster is currently down. If it can't be brought up, a failover will be issued

EDIT 00:00 UTC: Cluster is now available again, no failover happened.

Fixed · Services Logs · Global

Logs are currently partially unavailable through the console or CLI. Logs are still collected but display might not show current logs. They may also be out of order.

EDIT 06:22 UTC: Logs are now available again. No logs should have been lost but they might be out of order until 06:15 UTC.

Fixed · Cellar · Global

We are currently seeing elevated error rates on the old Cellar cluster. A few nodes went down making operations longer than usual, leading to timeouts or 500 / 503 errors. Nodes are already getting back up.

EDIT 00:21 UTC: The cluster is getting back to normal, errors have already significantly decreased and most of the requests should now be successful. We keep monitoring failed requests.

EDIT 03:00 UTC: No more failed request over the last 30 minutes, the incident is closed. We are still in the process of migrating this cluster data to the new cluster. Until we automatically migrate your buckets, you can migrate them yourself. Feel free to contact our support for more information

Fixed · Cellar · Global

The new Cellar cluster (cellar-c2.services.clever-cloud.com) had a brief interruption between 19:31:30 and 19:33:20 UTC on 23/08/2019 where most of the requests couldn't be handled or were dropped if already started. The problem has been identified and automatic actions have restored access to the cluster. The main issue will be investigated.

Deployments delayed
Fixed · Deployments · Global

Deployments are delayed, we are looking into it.

12:11: An orchestrator was experiencing intermittent network issues. The issue is now fixed.

Fixed · Global

From 20:30 UTC to 21:30 UTC, 16% of the hypervisors of the Paris zone failed to resolve the monitoring service domain name.

Applications which had instances on these hypervisors have been redeployed automatically because the monitoring could not reach them (even though they were available).

Fixed · MySQL shared cluster · Global

The MySQL c4 shared cluster of EU zone is experiencing issues. We are investigations.

EDIT 9:29 UTC: fixed.

July 2019

Metrics are unavailable
Fixed · Access Logs · Global

Metrics are unavailable, we are looking into it. Write requests are still processed.

13:27 UTC: Issue fixed.

Fixed · Infrastructure · Global

One of our hypervisor is experiencing issues and is unresponsive, we are restarting it. The applications on it have been redeployed on other hypervisors. Addons will be down during the restart.

EDIT 22:05UTC: the hypervisor is restarted.

EDIT 22:20UTC: incident fixed.

API maintenance
Fixed · Global

Our API will go under maintenance at 19:40 UTC. Deployments will be disabled for a few minutes and the Dashboard won't be available either.

EDIT 19:43 UTC: The maintenance is starting, API will be shortly unavailable.

EDIT 19:49 UTC: The maintenance is over!

Network Maintenance
Fixed · Global

An exceptional Network Maintenance is planned today at 20:00 UTC. Network interruptions are expected to happen from time to time during a few hours. They shouldn't last long. All applications will be redeployed on non-impacted servers, some add-ons will be unreachable at some point. Unfortunately, we couldn't postpone it due to calendar issues. Do not hesitate to ping us on the support if you have any questions.

EDIT 20:03 UTC: The maintenance should start shortly. We will keep you updated on its progress.

EDIT 20:53 UTC: The maintenance is still ongoing. Nothing unusual to report as of now

EDIT 21:20 UTC: Everything is going smoothly as seen in our tests. Nothing unusual to report as of now

EDIT 21:42 UTC: The maintenance is over. No network interruptions have been noticed by our monitoring systems. Everything is back to normal.

Fixed · API · Global

Our payment processor currently has troubles leading our calls to their API to sometimes fail. Multiple endpoints on our API request our payment processor's API and some of them will fail.

Here is a non exhaustive list of affected actions (some of them will succeed):

  • Application or add-ons creation
  • invoices payment
  • credit cards management

EDIT 23:30 UTC: Our payment processor issues should now be resolved. Everything should be back to normal on our side too.

Main API unavailability
Fixed · Global

A maintenance on components used by the main API will take place on 2019-07-11 at 10:00 UTC (12:00 CEST). The main API will be unavailable for a few minutes (up to 20).

10:02 UTC: Deployments queued now will be post-poned until the end of the maintenance.

10:04 UTC: The main API is now unavailable.

10:06 UTC: The main API is restarting.

10:09 UTC: Maintenance is over. The main API is available, pending deployments are starting.