Incident History
Full history of incidents.
October 2021
The Metrics / Access logs platform is currently having issues. We are investigating.
EDIT 11:00 UTC: A node from the cluster failed to reboot and was stuck in failed state. We are rebuilding this node. It will take 2 to 3 hours. No data will be lost.
(Times in UTC) 09:15 - The RabbitMQ cluster handling live logs started to fail with the "logs" vhost. We start creating the vhost again. 2021-10-24 08:00 - We notice that parts of the logs system are still not working. We investigate them. The Logs API keeps crashing for no apparent reason.
11:45 - The Logs API stopped crashing. We don't know why and continue to investigate the reason to fix this for the long term.
Webhook and e-mail notifications have not been sent since 22:30 UTC on 2021-10-21. The notification service lost its connection to the message queue service and failed to reconnect automatically. This was due to a short network outage between our two Paris datacenters. This issue has been mixed in with others and left unnoticed.
At 11:12 UTC today, the queue has been emptied so all webhooks matching the events during this period have not and will not be sent out. Events from 11:12 to 12:25 UTC have all been sent at once and everything is back to normal since then.
An add-on reverse proxy was unreachable on the PAR zone. Some applications might have had issues connecting to their add-ons or may have unexpectedly lost their connections to them.
The reverse proxy has been rebooted and this incident is now over.
One of our hypervisors needs to be shut down because of a faulty memory module. Applications have already been redeployed elsewhere and add-ons will be automatically migrated starting on October 19, 2021 at 20:30 UTC+2.
Impacted users will shortly receive an email and can contact us for any further questions.
EDIT 19/10 18:35 UTC: Migration of add-ons has started
There is an issue with the certificate associated with the *.cleverapps.io domains, which has expired. We are renewing it ASAP.
10:40 UTC+2: The issue is resolved.
The Metrics / Access logs platform is currently having issues, queries are returning errors. We are investigating.
EDIT 15:07 UTC: The problem has been identified and fixed. Queries should now be back, current data lag is 1 hour and 30 minutes. It should quickly come down in the next hour.
EDIT 17:58 UTC: Ingestion lag is now resolved
We are experiencing networking issues with our OVH-based infrastructure, we are looking for more information from OVH.
https://twitter.com/ovh_status/status/1448185498812485633?s=20
The website travaux.ovh.com is unreachable preventing us from getting a status on the maintenance where "No impact" was expected.
09:55 UTC+2: We still have no update from OVH.
10:01 UTC+2: https://twitter.com/olesovhcom/status/1448196879020433409?s=20
10:20 UTC+2: Our Montreal zone is reachable, others zones might come back soon.
All our zones are now reachable, you might still experience DNS issues or other issues due to the OVH incident it self.
There is an issue with the certificate associated with the cellar of our hds zone, we are investigating.
12:58 UTC: The issue is resolved.
Here is what we know so far:
The revocation server of the Certification Authority providing this certificate says that this certificate has been revoked on 2021-06-23, except it was still accepted just fine a few hours ago.
We have asked for a reissue of the certificate (this is an automatic operation). The reissued certificate has been installed and is working fine. Meanwhile, we have asked the CA about this revocation without any warning or notice and are waiting for an answer.
We are currently seeing git push errors at least when using the HTTP protocol, the connection gets refused. This mostly impacts pushes from our CLI.
We are investigating the issue.
EDIT 10:00 UTC: The issue has been fixed, pushes using the HTTP protocol should now be working as intended. Pushes and clones using SSH protocol were not impacted. We'll investigate further the issue.
Some applications have the wrong CC_JAVA_VERSION environment variable value, this may lead to unexpected deployment errors or runtime errors if the application redeploys. We are looking into it.
EDIT 18:45 UTC: CC_JAVA_VERSION should now be fixed with the right value. Impacted applications are redeploying to make sure they use the right version.
EDIT 18:58 UTC: If you changed the value of CC_JAVA_VERSION between 09:30 UTC and 18:45 UTC, the value might have been replaced with its previous version. Make sure you set it back to the right version if needed. Sorry for the inconvenience.
Our documentation is currently broken, we are looking into it. If you have any technical questions, feel free to contact our support team in the meantime.
EDIT 14:57 UTC: The problem has been fixed, the documentation should now be fully accessible at https://www.clever-cloud.com/doc/
A Let's Encrypt root certificate has expired yesterday. This can lead to various errors TLS error for old clients, the most common being "Certificate date is invalid".
Our Let's Encrypt certificates already provide the up-to-date Let's Encrypt chain but some older clients might not be able to trust that new chain because they don't have the new root Certificate Authority in their trustore. If you are in this situation with clients you can't update, we can sell certificates that will be trusted by those older clients. You can contact us on the support with the domains you need to protect.
You can also find more information about this expiration on Let's Encrypt website: https://letsencrypt.org/docs/dst-root-ca-x3-expiration-september-2021/
September 2021
Some API endpoints are returning high rates of HTTP 500 - Internal Server errors. We are investigating.
EDIT 14:57 UTC: A fix has been pushed, the errors should be resolved. We continue to monitor the situation.
EDIT 15:19 UTC: No more Internal server errors are happening, this incident is now closed.
The metrics and access logs are currently unavailable. We are looking into it.
EDIT 06:43 UTC: queries should be back to normal, the ingestion lag should take a few minutes to be consumed.
EDIT 11:12 UTC: Everything is back to normal
FSBuckets are not mounting properly on new deployments, we are investigating
Edit: New hypervisors were added but they had no support for fsbuckets yet.
Metrics and access logs are currently partially unavailable to query. We are investigating.
EDIT 14:35 UTC: The root cause has been identified, the ingestion lag currently sits at around 2 hours so metrics queries will be out of sync for the time being. Access logs are not ingesting and are currently kept in a separate queue. We expect the lag to start decreasing later tonight. This incident is a follow-up to the urgent maintenance of yesterday which mainly aimed at better stabilizing the cluster.
EDIT 23:34 UTC: Metrics have been fully ingested, access logs are still delayed but they are currently being written. Queries might still be slow, this is expected.
EDIT 6:30 UTC: The situation is back to normal.
Following maintenance, access logs and metrics may be unavailable for some queries as well as ingestion lag. Overviews of applications/organizations requests may also be impacted.
EDIT 16:17 UTC: the maintenance is still ongoing. Reads and writes are disabled since 15:42, this is expected.
EDIT 21:50 UTC: the maintenance is finished. Ingestion is catching up.
Logs & logs drains are experiencing issues.
EDIT 15:14 UTC: fixed, the related drains are currently catching up.
August 2021
Between 15:48 UTC+2 and 15:50 UTC+2, one of our reverse proxies on Paris unexpectedly timed out during a maintenance upgrade on our Paris zone.
The time outs last for about 2 minutes before the proxy was put out of the pool.
Some requests might have failed during the first minute and then, all requests handled failed during the remaining minute. Additional investigation will be performed to analyze what happened.