Incident History
Full history of incidents.
March 2023
The monitoring has detected that the metrics storage layer is offline. We are investagating.
EDIT 13:56 UTC : A node has crashed, the metrics storage layer finished its recovery process, it will take 20 minutes to consume the lag.
EDIT 14:22 UTC : Lag has been consumed and metrics storage layer has been operating normally
The monitoring has detected that a hypervisor is not responding. We are investigating.
EDIT 8:31 UTC: hypervisor is up and running.
We are currently investigating reverse proxies instabilities on our Paris zone.
EDIT 18:56 UTC: To be more specific about the instabilities, the connections were slower to be processed, increasing the response time, sometimes drastically. The root cause has been found and fixed at 18:42 UTC. Since then, everything is back to normal. We continue to monitor the situation.
EDIT 19:11 UTC: Additional investigation will be performed to pinpoint the exact cause of the problem and measures will be added to prevent it from happening again. Sorry for the inconvenience.
The main Clever Cloud API will go under maintenance for about 30 minutes, starting at 21:00 UTC.
During these 30 minutes, some deployments may not go through. Some calls may fail.
Everything seems to have gone well. The operation was over at 21:28.
EDIT 23:15 UTC: It seems like some application creation are having issues following this change, we are investigating.
EDIT 00:10 UTC: A fix has been implemented and applications are now correctly created. Some users may have had the API answer a 200 - OK for application creation but following requests for that application would return a 404 - Not Found. Sorry for the inconvenience.
(All times UTC)
- At 20:10 one of the 4 reverse proxies on zone PAR stops responding to some requests. No internal metrics changed, no weird logs were written. The requests would just time out. The other three were still running, so the requests errors were random.
- At 20:25 it stops responding at all.
- At 20:40 our external monitoring tool alerts us. We investigate, find which reverse proxy failed, restarted it.
- At 20:43 the reverse proxy is restarted and traffic goes fine.
(All times in UTC)
16:30 we started seeing alerts about high load on the primary node. 17:00 we started getting report about the cluster being unreachable. 18:00 after checking the cluster, we decided to restart the primary node.
Data may have been lost as the node was not writing / replicating correctly. We are still waiting for the primary node to restart. The secondary does not seem to elect itself as primary.
19:30 the secondary finally got promoted as primary. We are blocking users with unfair use of the cluster. 22:45 we detect that the node we restarted failed to get back in the cluster. We decide to remove it entirely and re-create that node from scratch. 2023-03-13 10:00 the node has fully reached the "SECONDARY" state. We put it back into production.
Measures have been taken to prevent future unfair use from users.
(All times in UTC)
11:30 Our main API keeps stopping to respond. We are investigating it. This impacts the following, in an irregular fashion:
clever sshmay not succeed- Some deployments may not go through
Applications should keep running, but some monitoring deployments may fail.
12:55 The API seems to have stabilized. The database seems to have had a huge load. We are investigating the queries responsible for that load and try to improve them.
We are currently investigating network issues on our Paris zone.
EDIT 17:15 UTC: The issue is now resolved. A part of our infrastructure in Paris couldn't access some public DNS servers anymore, leading to multiple DNS queries failing. An upstream network provider made a change that fixed the problem around 16:52 UTC.
Clever Cloud Core API is currently experiencing performance issues. We are investigating it.
EDIT 16:03 UTC: We are seeing improvements, we continue to monitor the situation and keep investigating the root cause. We continue to add more data collection around the various points of contention.
An hypervisor went down, we are investigating. Applications are being redeployed.
Update 11:11 AM UTC: The hypervisor has been rebooted, add-ons should be reachable. Root cause of the issue will be determined later. In the meantime, applications hosted on that hypervisor are still redeploying. We continue to monitor the situation.
Update 03:13 PM UTC: the same hypervisor went down again. It has been rebooted. Add-ons should be reachable. In the meantime, applications hosted on that hypervisor are still redeploying. We continue to monitor the situation.
Clever Cloud Core API is currently experiencing performance issues. We are investigating it.
EDIT 14:37 UTC: We are seeing improvements, we continue to monitor the situation.
EDIT 16:23 UTC: The incident is now over.
We are facing a network issue between MTL and our control plane causing some deployment issues. A workaround has been found and deployments are, as of now, OK on this region. A ticket has been opened in our subcontractor to solve the root cause.
EDIT 03/03 02:15 PM UTC: Connectivity between MTL and our control is to fully restored.
We are experiencing failures when deploying apps to RBX. We are investigating.
EDIT 10:32 AM UTC: a connectivity issue have been detected between RBX and our control-plane. The issue is now fixed.
February 2023
A maintenance has been planned on our Ticket Center tool February 28th, 2023 at 19:00 UTC. Users will need to refresh their Clever Cloud Console (https://console.clever-cloud.com) to complete the update. Otherwise, the Ticket Center might display an authentication error. During that time, actions on tickets (creation, comment, ..) might fail.
The maintenance is expected to last 5 minutes. If you urgently need to contact us, you can send an email to support@clever-cloud.com
EDIT 19:38 UTC: The maintenance is now over. Actions on the ticket center should be fully available. If you encoutner any problems following this update, please email us at support@clever-cloud.com
We need to conduct an update on our Jeddah hypervisors on February 28th, 2023. Services of impacted users will be migrated starting at 20:00 UTC before the update begins.
Impacted users will receive an email for each impacted service.
EDIT 2023-02-28 20:25 UTC: The maintenance is starting
EDIT 2023-02-28 22:18 UTC: The maintenance is now over.
We are currently experiencing degraded performances towards github.com services from our Paris infrastructure. We are investigating the issue. Tools relying on GitHub (composer, go, ...) might take longer than usual to fetch their dependencies or experience connections timeouts / instabilities.
EDIT 15:48 UTC: We are seeing improvements and the situation is currently back to normal. The root cause seemed to be a BGP announce change from GitHub's side that made our traffic go through suboptimal routes, leading to degraded performances. We keep monitoring the situation.
EDIT 16:30 UTC: The incident is fully resolved.
Clever Cloud Core API is currently experiencing performance issues. We are investigating it.
This is a follow up for the various hypervisors incidents we had those last weeks. A first batch of hypervisors will be updated to try and fix the issue. Impacted users will shortly be contacted by email.
The reboot is planned tonight (15/02/2023) at 22:00 UTC. Maintenance will start at 21:00 UTC.
EDIT 21:07 UTC: The maintenance is starting. Add-ons will be automatically migrated in the next few minutes.
EDIT 22:52 UTC: The maintenance is over.
An hypervisor went down, we are investigating. Applications are being redeployed.
EDIT 22:47 UTC: The hypervisor is back online with add-ons UP since a few minutes. Root cause of the issue will be determined later. In the meantime, applications hosted on that hypervisor are still redeploying. We continue to monitor the situation.
EDIT 23:44 UTC: The incident is now over. Sorry for the inconvenience.
We are currently seeing applications having troubles complete their deployments, especially when using dedicated build VM. They may be stuck or very slow at the cache archives upload. We are investigating.
EDIT 10:55 UTC: The root cause has been found. It was only impacting multipart uploads. For deployments already at the upload phase, you will need to cancel the current deployment and start a new one for the problem to be fixed. Sorry for the inconvenience.