Incident History
Full history of incidents.
February 2023
An hypervisor went down, we are investigating.
EDIT 22:24 UTC: The hypervisor is up again since 10 minutes. Add-ons are available again. We make sure all applications were redeployed.
EDIT 00:17 UTC: The incident is over.
At 16:29 UTC, a staff member started investigating an alert on one of our hypervisors. They saw the hypervisor could not be logged into anymore.
All services running on that hypervisor are still up and running, but deployments fail to stop the obsolete VMs and we cannot connect to the host itself. We are considering a "semi" kernel crash on the hypervisor's host. We are investigating and may reboot the hypervisor in the following minutes/hours. (First, we try migrating as much important services as possible to avoid causing too much downtime to our customers.)
EDIT 16:46 UTC: We are starting to migrate add-ons on the impacted hypervisor.
EDIT 18:54 UTC: We rebooted the hypervisor, everything went well, all the remaining services are UP again.
Between 12:39 UTC and 20:10 UTC, some users may have experienced an error message WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! when pushing code using git+ssh on our Git repositories. This was due to an update of the allowed signature algorithms of our SSH servers. Users that had an old signature algorithm stored in their known_hosts ssh file were impacted.
The change has been rolled back.
January 2023
A few customer complains about performance issues on MySQL shared cluster. We are investigating.
EDIT 10:00 UTC We have made a hardware upgrade to the MySQL shared cluster
Monitoring detect an increasing number of unreachable virtual machines. It seems related to an update deployment.
EDIT 01:00 UTC the update deployment has been rollback
Monitoring report that the number of timeout increase on the Clever Cloud API. We are investigating why.
EDIT 9:08 UTC : Backends behind Clever Cloud API are up and running. Numbers of timeouts have decreased. Everything is operating normally.
One server that host FSBucket need additionnal disk space.
EDIT 10:10 UTC Operation to increase the disk space is done. We are redeploying the associated applications
Live logs system has an issue with the storage backend that put it to read only mode
EDIT 22:56 UTC : The storage backend has left the read-only mode
Two hypervisors have rebooted in the Paris zone. Deployments have been impacted and some applications and databases may be unreachable. We are investigating the issues.
** EDIT 13:59 UTC ** One hypervisor is up and running
** EDIT 14:52 UTC ** The second hypervisor is down due to hardware issues
** EDIT 15:22 UTC ** Applications and databases may be difficult to reach as a load balancer node is hosted on the down hypervisor
** EDIT 17:00 UTC ** Deployments may have been impacted, we are redeploying the system
** EDIT 17:30 UTC ** Hypervisor is up and running. We are cleaning up the last thing
** EDIT 18:17 UTC ** Hypervisors are up and running. All systems seems working normaly
At 02:20 UTC, we started having alerts saying the Ceph pools are full. We are investigating this.
04:40 UTC, we take the decision to lower the replication ratio to let the cluster breathe.
A lot of backups failed, though. We will start them again during the day.
Several applications deployed on Paris region are not reachable. We are on it.
EDIT 21:56 UTC: we are experiencing a network connectivity issue, impacting parts of Paris region. Cellar is also impacted.
EDIT 22:01 UTC: Network connectivity is back online. Apps should be reachable. Cellar is in recovery, we are working on it.
EDIT 22:25 UTC: Cellar should be accessible. You may experience a bit more latency due to recovery processes in progress.
EDIT 22:57 UTC: Everything should be up.
An hypervisor in Paris is down/unreachable. We are investigating it. You may experience some deployments issues.
Edit 3:25 pm UTC: The hypervisor is back online. All impacted applications have been redeployed. If you are experiencing an issue, please contact our support.
We are currently facing an unavailability of the Metrics and access logs stack. The problem has been identified and we are working to bring it up.
Metrics through the console or Grafana or access logs query is currently affected.
EDIT 16:44 UTC: The service is back up, we are starting to process the backlog of events. You should now be able to query the data but it might lag a bit.
EDIT 17:01 UTC: The queue has been ingested. The service is now back to normal. Sorry for the inconvenience
Deployments are currently slower than usual. They may take more time to start or complete. We are investigating.
EDIT 13:46 UTC: The slowness is now resolved since 13:35. The initial cause of the slowness has been found and we continue to monitor the situation.
Deployments are currently unavailable and failing for unknown reasons. We are currently investigating.
EDIT 15:25 UTC: Deployments are running again. Some more operations will be done in the next few minutes to stabilize the situation. In the meantime, we continue to monitor the health of the deployment system.
EDIT 15:45 UTC: The incident is now over. If you still have troubles deploying your application, please reach out to our support team. Sorry for the inconvenience.
Cellar C2 is having issue with time sync. It may result with a "ClockTooSkewed" error when you try to list or access files.
We are working on fixing the clocks on the Ceph monitoring servers. (Ceph is the software we use to provide the Cellar service.)
EDIT 12:40 UTC+1: One of the reverse proxies in front of the Cellar system was desynchronized. This proxy is now out of the pool for further investigation and the issue should now be fixed.
We are currently experiencing network instabilities on the Paris zone. Our network provider is aware of the issue and we are currently awaiting for more information. One instability was detected at 9:42 UTC+1. No other since then.
Our metrics and access logs stack is currently unavailable, we are working towards bringing it back up.
Update 9:55 am UTC: Metrics and access log storage is now up. We are catching up the lag
Update 14:33 UTC: The lag of the Metrics and access logs platform is now resolved. Regarding the network instabilities, our network provider identified the issue and is working towards resolving it. It may take a few hours to get back to a nominal situation. We did not see any other instabilities since this morning.
Update 15:59 UTC: Another network issue happened at 15:50 UTC and lasted for ~1 minute, parts of the Paris zone was unreachable during that time.
Update 23:11 UTC: No other incident has been seen, we are still waiting for our network provider to ensure that the issue is resolved on their end.
Update 2023-01-09 14:18 UTC: We've seen two new events, one at 13:23 UTC and another at 14:14 UTC. We notified our network provider. Those may be related to the same problems we've seen last week.
Update 2023-01-09 19:47 UTC: Those two events weren't linked to the ones seen last weeks. The reason has been identified by the network provider and has been fixed. We are still waiting for confirmation of resolve on the original issue.
Some part of the authentication process in the API is failing. We identified the error and are trying to fix it.
December 2022
The Git repositories servers on all of our zones will go under maintenance today at 18:00 UTC+1. This maintenance will have the following impacts:
-
Delayed git repository creation for newly created applications
-
Delayed add or removal of SSH keys authorized to interact with the git repositories
GitHub applications will not be impacted.
During the maintenance, you will be able to continue to push your updates as well as do deployments. The maintenance is expected to last up to 1 hour. If you have any questions, please reach out to our support team.
EDIT 18:01 UTC+1: The maintenance is starting.
EDIT 18:35 UTC+1: The maintenance is now over. Thanks for your patience.
Due to an issue with our core APIs, we are doing urgent maintenance. It should take 15 minutes. Deployments will be blocked during this time. Applications will keep running.
EDIT 13:45 UTC - done.