Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

August 2022

Fixed · Infrastructure · Global

One hypervisor went down in MTL2. We are trying to reboot it.

It affects: 1 load balancer 1 redis add-on 1 mysql add-on The free postgresql databases on MTL.

Update 16:40 after investigating, we decide to redirect the IP of the load balancer to the second LB. A ticket is open at OVHCloud to investigate what seems to be a hardware issue. Update 17:56 OVHCloud team physically checked the server: the RAID card was broken. They changed it and restarted the server. Update 18:05 All VMs on the hypervisor are up and running again.

Fixed · Infrastructure · Global

An hypervisor on the Paris zone is currently unreachable. We are looking into it.

EDIT 17:38 UTC: Hypervisor has been rebooted. Services are being restarted.

EDIT 18:08 UTC: Services have all been restarted. We continue looking into why the hypervisor went down and continue to monitor the situation.

EDIT 18:27 UTC: Initial investigation shows that a KVM kernel bug was encountered, leading to a kernel crash. We will investigate further to see if this can be mitigated by an update. The incident is now over.

[New York] Network loss
Fixed · Infrastructure · Global

We are seeing a network loss towards the New York zone from multiple places since 06:05 UTC. We are looking into the issue. Applications and add-ons may not be reachable from different places and multiple services on the zone (deployments, logs) will not be available.

EDIT 07:04 UTC: We are seeing network improvements to reach the zone. It is currently operational but we are still waiting on confirmation from our provider. From our point of view as of now, traffic towards the zone was dropped when reaching the Level3 network transit. Our network provider seems to have changed it to another provider, allowing us to reach the zone again.

EDIT 12:18 UTC. The network problem is fully resolved. We are still waiting for an incident report from the network operator of the Datacenter. We will share it once available.

EDIT 2022-08-26 14:27 UTC: Here is the report from our provider: It has been identified that the incident is due to a bug found in our device at DRT1. As an initial resolution, our team rebooted the device. Consequently, all alarms cleared and all services were restored after executing the said activity. As of the moment, we can confirm that the link has remained clean and error-free since the service went up.

Fixed · Infrastructure · Global

We are investigating an unresponsive hypervisor on the Paris zone. An FSBucket server is on this hypervisor, some PHP applications may be impacted as well as add-ons hosted on this hypervisor.

EDIT 18:02 UTC+2: Hypervisor is rebooting

EDIT 18:04 UTC+2: Hypervisor is up again. Services are currently restarting.

EDIT 18:25 UTC+2: Hypervisor services are all up since a few minutes. Add-ons should now be reachable. Applications of owners using the FSBucket server that is hosted on this hypervisor will be redeployed. Since there is a huge number of applications, you can deploy them on your end directly if needed. We will continue to monitor the situation.

EDIT 19:10 UTC+2: The situation seems to be back to normal. We will investigate further why this hypervisor became unresponsive. If you still have any issues, please contact our support team.

Lost FRS server
Fixed · FS Buckets · Global

We lost a server hosting FS buckets. server up and running

Fixed · FS Buckets · Global

A FS Bucket machine has become unreachable from 2:59 PM to 3:01 PM today. It has been rebooted and is now available.

Fixed · Reverse Proxies · Global

An reverse-proxy for add-ons has become unreachable from 3:53 PM to 3:59 PM. It was rebooted.

July 2022

Deployments slow down
Fixed · Deployments · Global

We are experiencing slow down of deployment, we have identified the root cause and are working on the solution.

EDIT: 00:54 Issue has been resolved, deployments must be worked normally

Fixed · Global

At 19:00 UTC+2 on Tuesday 26th July 2022, our tickets center the Clever Cloud Console will be unavailable for a few minutes.

Once the maintenance is over, you will have to refresh your Clever Cloud Console to be able to access your tickets or contact our team.

During this maintenance, you will still be able to reach our support team using our email address: support@clever-cloud.com

EDIT 2022-07-26 18:59 UTC+2: The maintenance is about to start.

EDIT 2022-07-26 19:10 UTC+2: The maintenance is now over. You will need to refresh your Clever Cloud Console to access the ticket center.

Fixed · RabbitMQ shared cluster · Global

The cluster refused publishers messages since 12:32:34 UTC due to a system wide alert that stopped all nodes from accepting publishers messages. This means that RabbitMQ clients would keep trying to publish their messages until the cluster accepted them. The node has been restarted at 17:23 UTC, fixing the issue.

Investigations will be carried out to understand how this happened and why our monitoring did not raise an alert.

The cluster should now be fully operational.

One FSBucket node is down
Fixed · FS Buckets · Global

One of our FSBucket node is currently down. We are working to resolve the issue.

EDIT 10:40 UTC - fixed.

Deployments may be blocked
Fixed · Deployments · Global

In rare occasions, an inappropriate behavior of our scheduling infrastructure can lead to deployment being stuck. We've identified the root cause and we're qualifying a fix. If it happens, don't hesitate to reach our support team.

Fixed · Deployments · Global

We are experiencing connectivity issue between NYC and PAR. These connectivity issue are impacting deployments on the NYC zone. We are working on it.

EDIT 16:03 pm: connectivity has been resolved

Metrics maintenance
Fixed · Access Logs · Global

In our efforts to stabilize the Metrics infrastructure, we will perform a maintenance on 13 of July. Once it is started, some lag can be expected for a few hours.

Maintenance will start at 07:30 am UTC

EDIT 07:30 am UTC: Starting maintenance

EDIT 08:16 am UTC: Maintenance is over, we are catching up with the lag

EDIT 08:30 am UTC: Queries are currently disabled to speed up recovery

EDIT 09:17 am UTC: our maintenance triggered a major compaction on our storage layer. To speed up recovery, query are still disabled

EDIT 16:20 pm UTC: major compaction is over. We are struggling to handle both read and write operations at the same time. We are working on it.

EDIT 20:23 pm UTC: queries are still disabled. We are testing new configurations to resolve the issue

EDIT 14 of July 9:22 am UTC: it's a brand new day, we are still working on it.

EDIT 14 of July 18:26 pm UTC: We are struggling to handle both read and write operations at the same time. We are working on it. Happy french national day.

EDIT 16 of July 17:35 pm UTC: We found a performance issue triggered when the dotmap on the Console is accessed. We disabled some macros used to retrieve data to allow other users to access metrics. Metrics and access logs are now accessible.

Fixed · Cellar · Global

The service was having troubles handling most of the requests between 11:24 and 11:28 UTC. We will investigate further the issue. The Cellar service is currently operational.

Fixed · Infrastructure · Global

Starting 09:31 UTC, we saw intermittent network failures on the Roubaix (RBX) zone hosted on OVH. Failures are both from the external and internal networks. Timeouts reaching your applications or add-ons might have happened.

Some applications are being redeployed for Monitoring/Unreachable because the monitoring couldn't see them anymore.

Things seem to be working fine again since 09:37 UTC. We continue to monitor the situation and will try to get more information from OVH.

EDIT 11:12 UTC: The issue has not occurred again. We will wait for any input from OVH and will add it here if we get any useful information.

Fixed · Infrastructure · Global

An hypervisor has been lost on the OVH Roubaix zone. We are investigating. Impacted services are FSBuckets and add-ons.

EDIT 15:32:00 UTC: The server is back online. We are making sure services are correctly restarted. Additional services were impacted: One application reverse proxy and one add-on reverse proxy were unavailable.

EDIT 15:48:00 UTC: We are still investigating the cause of the reboot. We opened a ticket on OVH services to know if they had any un-planned intervention for that machine.

EDIT 16:03:00 UTC: The machine is unreachable again. We are investigating.

EDIT 16:11:00 UTC: The machine is up again. We are starting to suspect a hardware issue.

EDIT 16:30:00 UTC: We will drop all services from the machine to avoid any other issues until we know more about the underlying issue. FSBuckets server will be moved out around 19:00 UTC.

EDIT 19:59:00 UTC: Unfortunately, FSBuckets are going to require more time to move to another server. So far the server is working fine but OVH suspects an issue with the power supply.

EDIT 23:58:00 UTC: The FSBuckets migration is starting. FSBuckets will be set into read-only and applications will be redeployed to use the new server.

EDIT 2022-07-09 00:28:00 UTC: Buckets are fully migrated. The server is now empty and will be investigated further by OVH. This incident is now over.

[PAR] Network maintenance
Fixed · Global

A network maintenance has been scheduled by our network provider for Wednesday 06/07/22 at 22:30 UTC. The maintenance should not have any visible impacts other than a few seconds of network delay while the network links switch to the backup links.

EDIT 22:30 UTC. The maintenance is starting.

EDIT 22:55 UTC: Maintenance is over, no visible impact happened, links failed over in less than 100ms each time.

Ingestion queue issue
Fixed · Access Logs · Global

One of the server queue storage reach its disk max storage capacity

One of the partition is corrupted, fixing

EDIT 17:10 UTC: The underlying issue has been fixed. The queue is currently being processed. Some events might have been lost during the cluster rebalance. Data points will take a few more hours to be up-to-date in the various dashboards.

EDIT: Queue is in sync

Fixed · API · Global

We are currently looking into it. Console and CLI are not working correctly.

A batch was sent by an employee. The throttle interval was set two small and the batch made a huge amount of queries to the database, making it unresponsive. We stopped the batch and will restart it with a higher throttle interval.