Incident History
Full history of incidents.
August 2024
(Times are in UTC)
- At 14:24, two of the add-ons reverse proxies of the PAR region stopped responding. After investigation, we found out that the two failed to reconfigure correctly, due to a "stucked" port: the port was considered still used and fail to switch between the old process and the new.
- At 14:34, we decided to fully reboot these two reverse proxies. It successfully fixed the issue.
The consequence of this incident is that some applications that were trying to use one of these two reverse proxies (of a total of 7 proxies) lost their connection to the database for 10 minutes.
The monitoring has detected errors on read queries of the telemetry cluster. We are investigating.
EDIT 21:30 UTC : We found out that the issue is related to indexes of the time series database, we are investigating the reason of the error.
EDIT 21:40 UTC : Some indexes had errors and have been rebooted, the estimate time to recover indexes is around 01:00 UTC.
EDIT 01:00 UTC : Indexes are still rebooting, the new estimate time is 03:00 UTC.
EDIT 02:47 UTC : Indexes are back online and query is available.
EDIT 07:30 UTC : We are running some maintenance operation, the query may be hanging a bit.
EDIT 08:00 UTC : We have shutdown the query to get some place to our maintenance query to run as fast as possible. We have found the root cause issue and we are fixing it, but to resolve read errors, we also need to achieve some clean up in parallel.
EDIT 09:40 UTC : We have turn on the query again, we have still maintenance queries running in the background.
EDIT 13:00 UTC : We have turn off the query, we are struggling the reads with the maintenance queries. To reduce the time of the recovery process, we took the decision to shutdown the read queries to keep the maximum compute space to the maintenance ones.
EDIT D+1 08:00 UTC : We have turn on the query again, the maintenance queries has finished during the night.
We are investigating the loss of an hypervisor on the MTL region.
EDIT 16:36 UTC+2: The machine seems to have an hardware problem. Our provider is investigating the issue.
EDIT 17:36 UTC+2: We've been informed that this server was concerned by this maintenance: https://network.status-ovhcloud.com/incidents/ldl56trpj3kk. We are looking at how much time they need to complete this maintenance.
EDIT 17:48 UTC+2: The hypervisor has been rebooted by OVH. We are currently checking its state and restarting services.
EDIT 18:03 UTC+2: The incident is now over.
We may be impacted by https://network.status-ovhcloud.com/incidents/nnhpfdw50vsn which we are investigating. Only OVH regions based services are concerned.
Update 13:33 UTC: we are indeed impacted by OVHcloud's Backbone incident. Some network routes cannot reach OVHcloud's datacenters. We are working on it. More info can be found on https://x.com/olesovhcom/status/1819742478586528146
Update 14:42 UTC: network seems more reliable now. We are still watching the network links
Update 15:16 UTC: The services are getting operational according to OVHcloud and we are not seeing network issues anymore.
Phone number for on duty call of some customer experience a problem in our provider of telecommunication subsystems. Phone rings, but after there is impossible to talk on the phone. Customers with a problem, need to send directly by mail support@clever-cloud.com And they will be called back.
We are experiencing a global outage. We observed a network split in addition to an event bus outage. The effect has been inpactful for some core services.
EDITS :
- 2:00 PM CEST - Core services are being recovered and Deployments are being reloaded. This will synchronize back load balancers for customer's application trying to reach their new deployments.
- 2:08 PM CEST - Some services are being shut to accelerate the recovery process. Expect disturbed experience for observability and deployments for a few minutes
- 2:29 PM CEST - Criticial Core services are OK. Deployments are being rolled out.
- 3:07 PM CEST - Some workload queues have still difficulties to be processed. Some components may still be in an unstable state. Current effort is to identify them, then reload them.
- 3:40 PM CEST - Some hypervisors have experienced some crashes. Recovery process is occuring and will take a couple of minutes
- 3:56 PM CEST - Some hypervisors seems still experiencing network issues.
- 4:16 PM CEST - Apps are being deployed for premium customers. All apps are going to be deployed. Anyone can accelerate the process for its own application by manually deploying them.
- 4:24 PM CEST - In the meantime, we continue to identify noisy VMs that have been impacted by the outage
- 5:15 PM CEST - Metrics API is being restarted.
- 6:20 PM CEST - Last deployments are being rolled out. Reminder : accelerate by triggering a redeploy action
- 6:30 PM CEST - Still a few hundreds of VMs are consuming very high CPU rates and being cleaned.
- 6:35 PM CEST - We estimate approximately 40min to have full recovered all deployment of applications (MANUALY REDEPLOY FOR FASTER RECOVERY)
- 7:05 PM CEST - All IPSec links should be back online
Following https://www.clevercloudstatus.com/incident/877, we have difficulties to process access logs, you may observe holes and lags.
Following https://www.clevercloudstatus.com/incident/877, some deployments are failing. We currently working on a solution.
EDIT: 10H31 UTC - A workaround has been found to ensure that deployments work again
Connections issues (producers/consumes) during cluster upgrade
It can lead to fail in app redeployement
Order of DEV add-on is currently locked. No impact on existing add-on.
We are investigating
[EDIT 12:00 CEST]: we have identied and fix the lock
Because of the hardware issue describe in https://www.clevercloudstatus.com/incident/874, we need to rebalance data on Cellar North. Customer may experience higher latency than usual.
An hypervisor on the GRA-HDS region is unreachable. We are working on it.
EDIT Thu Aug 01 09:13:09 2024 UTC: hypervisor has been rebooted. A hardware issue has been detected. All applications have been redeployed and there was no customer databases on the hypervisor.
July 2024
Application deployement take unbound time to proceed.
We are investigating the issue
09:30 UTC We notice deployment perturbation. 10:02 UTC We have found a cause of the perturbation and fixed it.
14:30 UTC We notice other issues with the deployment system. 16:30 UTC After further investigations, we found the cause of the perturbations. We applied a temporary fix. The deployments are back on tracks!
We are working on a stronger fix for the deployments.
We identified a bottleneck on our FoundationDB cluster for warp10-c2.
Writing is impacted and might occur a lag in read metrics. We enabled sampling on data.
6:32 UTC: We identified unusual usage that was harming the system
6:35 UTC: Unusual usage stopped, the storage layer is starting to recover
6:58 UTC: Storage layer fully recovered, we still investigate & watch over the system
7:45 UTC: System is back to normal
At 2024-07-12 23:35 UTC, we received an alert about WSW hosts not responding. We checked and coud not ping any of our servers.
At 23:43 We pinged again. A ssh connection to the hypervisors allowed us to see the servers had an uptime of 1 minute. We checked that all services running on the servers restarted correctly and fixed those that were not correctly running. Applications have been redeployed by the monitoring. At 23:55 everything seemed to be back to normal.
We don’t know yet why the servers were rebooted.
Some Pulsar brokers are having issues connecting to the underlying zookeeper. We are investigating the reason.
There was an issue with zookeeper sessions. It is now fixed.
We've updated load balancer IP addresses for applications and websites hosted on Clever Cloud. The new IP addresses now in use are:
91.208.207.214
91.208.207.215
91.208.207.216
91.208.207.217
91.208.207.218
91.208.207.220
91.208.207.221
91.208.207.222
91.208.207.223
Important:
We are going to remove 4 IPs that you must stop to use between now and August 23rd, 2024:
46.252.181.103
46.252.181.104
185.42.117.108
185.42.117.109
After this date, your applications and websites will no longer be able to use these IP addresses.
We still recommend to use CNAME DNS records when it's possible. To ensure that there is no disruption to your applications and websites, please make sure that your apex domain names are updated to point to the new IP addresses. You can update your apex domain names by editing the DNS records for your domain.
Impact:
There should be no downtime for your applications or websites as a result of this change. However, if you do not update your apex domain names before August 23rd, your applications and websites may be unavailable.
What you need to do:
Review your apex domain names and ensure that they are pointing to the new IP addresses. If you are unsure how to update your apex domain names, please contact your domain registrar or Clever Cloud support.
For more information:
Please refer to the Clever Cloud documentation for more information about load balancers and DNS records: https://developers.clever-cloud.com/doc/administrate/domain-names/#using-personal-domain-names
You can take a look at the changelog entry about this change: https://developers.clever-cloud.com/changelog/2024-06-28-new-ip-list-paris
You can also contact Clever Cloud support if you have any questions.
Time slot: 09/07/24 from 07:00 AM UTC to 09:00 AM UTC
Our infrastructure provider will perform hardware maintenance impacting our whole Warsaw region. As there is an electrical outage risk for servers, we will follow their advice to shut down the region during the maintenance that may last up to 2 hours (from 7:00 AM UTC to 09:00 AM UTC). In case your services cannot bear such unavailability, we advise you to migrated them to another region such as Clever Cloud Paris before the maintenance.
You can request assistance by reaching out to our support.
EDIT 2024-07-09 10:05 UTC: The maintenance has completed. All service are up and running on the WSW region.
Some logs drains are not correctly sent to their target. The issue has been identified and is being resolved.
EDIT 17:00 UTC+2: Drains should be available again. The incident is now over.
Addon logs are not available, there is an outage on the Elasticsearch cluster
08:23 es sink has been paused to restore logs Drains
system restored