Incident History
Full history of incidents.
October 2022
At 04:30 UTC: a pulsar cluster started to behave strangely (See https://www.clevercloudstatus.com/incident/574 ) At 05:30 UTC: on PAR, notification services on the hypervisors try to send messages in a loop, filling the system with stuck processes. At 07:00 UTC: the OS of these hypervisors start to kill processes to make room. It impacted some applications and databases. We start working on shutting down the stuck processes and restarting the broken instances. At 10:00 UTC: we finish restarting all the broken instances.
Between 10:25 UTC and 20:25 UTC, some applications hosted on the OVH RBXHDS zone may have experienced random 503 response errors due to faulty reverse proxies. The issue has been found and is now resolved.
Additional investigations will be conducted to understand why our monitoring system did not report the issue earlier. Apologies for the inconvenience.
Add-on APIs database cluster disk is nearful. We are migrating it to a bigger disk.
Operation will take 10 minutes, during which add-on API will be unreachable.
(Times are UTC) 04:45 - Deployments are broken because of a pulsar issue. We are investigating.
05:45 - To prevent issues on the infrastructure, we disabled all deployments.
05:55 - We detect that some VMs are DOWN. It seems that the pulsar connection issues have overwhelmed the hypervisor's processes.
06:05 - We shut down the processes that fill up the hypervisors. It seems to fix the issue.
06:20 - The deployments seem to be back on tracks. We continue investigating the pulsar issue before putting it back into the deployment processes.
09:09 - We are still experiencing deployments issues. We are investigating.
12:28 - Deployments have been fixed.
We are observing high latency on our reverse-proxies on PAR.
It looks like we are under a DDoS. We are monitoring it and blocking IPs that are performing the most requests.
EDIT 15:08 UTC: we have found the application that was taking 50% of all the platform traffic. We blocked all the IPs trying to reach that application. Traffic is now operational.
September 2022
Our distributed database responsible for metrics and access-logs storage is not ingesting fast enough. As a result, you may experience some lags during queries. We are investigating.
EDIT 16:06 UTC: Ingestion lag is now resolved.
Our "main" API is very slow. We are investigating to find out why.
Some components related to ingestion of metrics and access logs are currently overloaded. We are working on it.
** 16:30 UTC **: Incident has been resolved
Some Elasticsearch add-ons are currently reporting a license expiration. The license is set to expire on 2022-09-30 23:59:59 UTC. Our team is currently working on it and the license of affected add-ons will be updated prior to the expiration date. We will update this incident once all add-ons are updated.
No service degradation is to be expected from this warning.
Please reach out to our support team should you have any questions regarding this matter.
EDIT 2022-09-29 17:30 UTC: A first license update has been applied. A new license update will be applied in the following days to finish the license update.
EDIT 2022-10-12 16:55 UTC: All licenses have been updated with a valid platinum license. The incident is over.
The service encountered an outage in the ingestion path which retained the access logs at the messaging-level. While being identified by the monitoring, this error unfortunately triggered as low-priority, hence being silent during a part of the week-end which led to a drop of access logs after a retention period. We have found the root issue and the problem is now resolved and should prevent further similar incident. Besides, we've fixed the level of criticity in our monitoring infrastructure.
26/09/2022 12:00 UTC: End of incident
A maintenance will occur on the Grafana used to plot Clever Cloud metrics the 09/27/2022 at 2:30 p.m. (CEST). We will update our instances to the last major release of Grafana: Grafana 9. You can check Grafana release post to learn what this change will bring to you: https://grafana.com/docs/grafana/latest/whatsnew/whats-new-in-v9-0/.
Applications creation might fail for Jenkins runners with HTTP 500 Internal Server Error. A fix will be soon deployed to fix the underlying issue.
EDIT 16:46 UTC: The fix has been deployed. We are monitoring the situation. This issue also impacted Heptapod runners creation.
EDIT 17:28 UTC: The issue has been fixed, runners creation are now working correctly. Sorry for the troubles.
We are currently seeing network loss between our Paris infrastructure and our zones on OVH (Roubaix, Montreal, ...). We are currently investigating the issue.
EDIT 11:50 UTC: First investigations are showing that it is not only a network issue between our Paris infrastructure and the OVH network. It seems to impact other network links as well. We will reach to OVH and try to know more about it.
EDIT 11:51 UTC: The incident has been renamed from "Network issues between Paris and OVH zones" to "Network issues on OVH zones"
EDIT 11:58 UTC. We are seeing improvements since a few minutes now. Connectivity has been restored from our point of view. We keep waiting for more information.
EDIT 12:09 UTC: We have not seen any new disruption so far. We consider this incident closed while we wait for a more detailed incident report from OVH.
EDIT 12:59 UTC: OVH status: https://network.status-ovhcloud.com/incidents/5mldyhd6v99c
Some HBase datanode have lost their regions all datanodes are OK
An hypervisor needs to be rebooted on our Montreal zone. The reboot will happen at 08:00 UTC on Friday, 16 September. Add-ons that support automatic migration will be migrated automatically starting at 07:30 UTC. You can also perform the migration at a time that suits you more before the given deadline.
This will also impact some FSBuckets add-ons during which reads and writes will be unavailable. Applications will be redeployed automatically once the maintenance is over to make sure they correctly re-connect to the FSBucket server.
The maintenance is expected to last 15 minutes.
Impacted users will shortly receive an email with the impacted add-ons.
** Edit 08:05 UTC ** Waiting for last migration to end
** Edit 08:25 UTC ** Last migration has ended, the maintenance is beginning
** Edit 08:35 UTC ** The server has rebooted successfully
** Edit 08:55 UTC ** Everything is up and running normally
A file-system bucket server was down during 6 minutes beginning at 12:24 UTC and ending at 12:30 UTC on Paris data center.
We have fix the issue and watching the service.
Some tokens used by our infrastructure have not been renewed. As a result, some vms cannot push their latest metrics We are working on it.
EDIT 10:38 UTC: All expired tokens have been regenerated and updated. Sorry for the inconvenience.
Our distributed database responsible for metrics and access-logs storage is not ingesting fast enough. As a result, you may experience some lags during queries. We are investigating.
EDIT 03/09/2022 12:10 UTC: lag is finally catching up, we will keep you posted.
EDIT 03/09/2022 16:10 UT: lag is fully recovered
August 2022
An hypervisor needs to be rebooted on our Paris zone. The reboot will happen at 22:00 UTC on Monday, 5th September. Add-ons that support automatic migration will be migrated automatically starting at 21:00 UTC. You can also perform the migration at a time that suits you more before the given deadline.
This will also impact some FSBuckets add-ons during which reads and writes will be unavailable. Applications will be redeployed automatically once the maintenance is over to make sure they correctly re-connect to the FSBucket server.
The maintenance is expected to last 15 minutes.
Impacted users will shortly receive an email with the impacted add-ons.
EDIT 2022-09-05 21:10 UTC: Add-ons migrations is starting
EDIT 2022-09-05 21:40 UTC: Add-ons have been migrated. The hypervisor reboot will happen in twenty minutes.
EDIT 2022-09-05 22:00 UTC: Hypervisor is rebooting
EDIT 2022-09-05 22:28 UTC: Hypervisor has been rebooted in 4 minutes, fsbucket server went back one minute later with most clients reconnecting. We started all affected applications to make sure everyone properly reconnects.