Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

May 2022

Fixed · Infrastructure · Global

We are currently having an unreachable hypervisor on the Paris zone due to a connection loss. We are trying to restart it. Impacted applications are automatically redeployed.

EDIT 22:46 UTC: The hypervisor doesn't reboot, we continue our investigation.

EDIT 00:06 UTC: The hypervisor is back online since a few minutes. All services are now available again. The extended period of downtime has been identified and will be fixed on similar hypervisors to have a faster recovery next time.

Fixed · Access Logs · Global

Ingestion of new access logs and metrics points is currently having an issue, leading to missing data points in metrics. Access logs ingestion is currently on hold and will be processed later. The issue has been identified and we are working to fix it.

EDIT 21:04 UTC: Ingestion is now back to normal. Access logs will be processed over the next few hours.

[PAR] Server lost
Fixed · Infrastructure · Global

We lost a server which host severval components on PAR zone

UPDATE: all applications have been redeployed

Fixed · Reverse Proxies · Global

Some applications are experiencing issues. We are investigating it.

UPDATE 14:57 UTC: Some Add-ons are being inaccessible due to a faulty proxy. We're removing it from the pool to mitigate.

UPDATE 14:59 UTC: Services are being reloaded to ensure the faulty proxy is removed from the pool.

UPDATE 15:10 UTC: Services are back online for redeployed apps. A faulty sentry induced an abnormal behaviour in the API.

CALL FOR ACTION 15:23 UTC: Remaining applications are currently redeployed. If you're impacted, we advise you to redeploy your app to accelerate the recovery process

Fixed · Deployments · Global

We currently have issues with deployments. Deployments may end up with errors asking you to contact our support alongside a stacktrace. We are currently working on a fix.

EDIT 14:59 UTC - We have identified defaulting component which encounters an issue in the connection pooler.

EDIT 15:09 UTC - deployments queue is being consumed and catching up. Issue it mitigated.

EDIT 15:23 UTC - Incident is fixed.

Root cause: we've found an issue in a messaging driver on a couple of isolated servers. Anyway, we've curated out this specific driver to fall back on an alternative messaging layer. In the coming days, we will dive into this specific bug we've found and will communicate the bug fix upstream.

Fixed · MySQL shared cluster · Global

The MySQL c5 shared cluster is experiencing issues. We are investigating.

EDIT 20:02 UTC: the MySQL shared cluster is back online.

Fixed · Services Logs · Global

Logs are currently having some ingestion/query issues. We are working on it.

EDIT 21:39 UTC - querying logs is now available.

Fixed · MySQL shared cluster · Global

The MySQL c6 shared cluster of EU zone is experiencing issues. We are investigating.

EDIT 21:39 UTC - shared cluster is now back online

Fixed · SSH Gateway · Global

The SSH Gateway will undergo a maintenance which will stop the service. Expected downtime is 30 minutes. During this time, SSH access to instances will be unavailable both from the CLI or from the regular SSH tool. Existing SSH connections will be stopped.

Maintenance is expected to start in a few minutes

EDIT 17:56 UTC: Service is back online, you should now be able to SSH to your instances. Sorry for the inconvenience.

Fixed · Access Logs · Global

Metrics and access logs are currently having some ingestion/query issues. We are working on it.

EDIT 23:06 UTC - Storage cluster is now up. We are now catching up the accumulated ingestion lag. Query components will be restarted in a rolling fashion throughout the next 6 hours.

EDIT Sunday 11:27 UTC - Some query components are still reloading

EDIT Sunday 20:27 UTC - We are still experiencing issues on the query components.

EDIT Monday 07:20 UTC - Query is back online

Fixed · Infrastructure · Global

A few hypervisors on the Paris zone had a configuration issue between 12:21 UTC and 14:16 UTC leading to instances not being properly monitored. This caused Monitoring/Unreachable deployments for the instances hosted on them.

Because of this, those hypervisors became more empty than the others. More VMs were scheduled on them since they had more resources available, which then lead to more Monitoring/Unreachable events.

Instances weren't, for the most part, unreachable, but were redeployed anyway.

This should now be fixed. Sorry for the inconvenience

PAR: FS Bucket Migration
Fixed · Global

Some FS-Bucket add-ons will need to be migrated to a different server for security reasons. During this migration, the Buckets will be in Read-Only mode. Any attempt to create or update a file on the add-on will fail, including for FTP operations. Errors related to Read-only file system are expected during this migration.

The migration is expected to last at most 1 hour. All impacted applications will be redeployed during the migration. After the deployment, applications will be able to write to the bucket. Read operations will not be impacted.

Users of buckets that need to be migrated have received emails.

EDIT 2022-05-31 10:00 UTC: The migration is starting, buckets will be put into read-only.

EDIT 2022-05-31 10:25 UTC: The migration is over. Applications have started redeploying, it should take around 2 hours. You can redeploy your application earlier to finish the migration.

EDIT 2022-05-31 13:11 UTC: All applications have been redeployed, the migration is now over.

Fixed · Access Logs · Global

Metrics and access logs are currently having some query issues. We are working on it.

EDIT 07:16 UTC - Indexes have been rebuilt. Query is now available.

Fixed · Infrastructure · Global

There is currently a delay in monitoring actions for some applications. This may result in extended time to detect crashed application instances and upscales / downscales events. Actions are currently queued and will resume shortly. ETA is 30 minutes.

EDIT 17:12 UTC: The queue is still being consumed.

EDIT 17:27 UTC: The queue is now empty. Every monitoring actions should now be working as expected.

Fixed · Infrastructure · Global

An hypervisor is currently unavailable. Applications are currently restarting. Add-ons hosted on that hypervisor are currently unavailable. We are looking into the root cause.

EDIT 10:50 UTC: Hypervisor is back online. Add-ons hosted on that hypervisor are currently available.

Fixed · Access Logs · Global

Metrics and access logs are currently having some query issues. We are working on it.

EDIT 15:02 UTC - Indexes have been rebuilt. Query is now available.

Fixed · Access Logs · Global

Metrics and access logs are currently having some query issues. We are working on it.

EDIT 09:20 UTC - Indexes have been rebuilt. Query is now available.

Fixed · Infrastructure · Global

A FSBucket server was unreachable for 15 minutes, leading to increased response time for basic read / write operations on some FSBuckets. This has been fixed, impacted applications will be redeployed.

PAR: FS Bucket Migration
Fixed · Global

Some FS-Bucket add-ons will need to be migrated to a different server for security reasons. During this migration, the Buckets will be in Read-Only mode. Any attempt to create or update a file on the add-on will fail, including for FTP operations. Errors related to Read-only file system are expected during this migration.

The migration is expected to last at most 1 hour. All impacted applications will be redeployed during the migration. After the deployment, applications will be able to write to the bucket. Read operations will not be impacted.

Users of buckets that need to be migrated have received emails.

EDIT 24/05/2022 12:00 UTC+2: The migration will start soon. FSBuckets will be put into read-only for a couple of minutes so that all buckets are correctly synchronized.

EDIT 24/05/2022 12:03 UTC+2: FSBuckets are now in read-only mode.

EDIT 24/05/2022 12:39 UTC+2: Synchronization is over. Applications are being redeployed. If you wish to recover faster, you can trigger a deployment through the web Console or CLI. Deployments are expected to all be started within the next 30 minutes.

EDIT 24/05/2022 13:34 UTC+2: The migration is over, if you have any issues, please contact our support team

Fixed · Access Logs · Global

Metrics and access logs are currently having ingestion and query issues. We are working on it.

EDIT 08:28 UTC - We are consuming the lag.

EDIT 08:28 UTC - Indexes are rebuilding.

EDIT 09:34 UTC - Indexes are rebuilt. Query is available.

EDIT 16:03 UTC - Fixed.