Incident History
Full history of incidents.
March 2021
Around 50% of TLS connections made to one of the HTTP/2 reverse proxies were dropped indicating a lack of certificate. The issue's origin was a misconfiguration of this reverse proxy. Additional checks have been put in place to prevent this from happening again.
The error started at 12:23:36 UTC and stopped at 12:46:50 UTC, lasting around 23 minutes.
We are experiencing issues with our metrics/accesslogs storage cluster.
EDIT 13:21 UTC - fixed.
Metrics & AccessLogs querying components are temporarily unavailable.
EDIT 17:03 UTC - fixed.
Our Liar Proxy hosted on OVH is currently unavailable since 01:23 UTC+1. The incident on OVH side is http://travaux.ovh.net/?do=details&id=49473& but the Strasbourg (sbg) zone seems to be having a more general issue: http://travaux.ovh.net/?do=details&id=49471
We'll update this incident in the morning. Until then, if OVH fixes the issue before that, the liar proxy should recover network access.
EDIT 11:06 UTC+1: This service is in SBG1 which is currently impacted by the fire that took place in SBG. It may take several days to come back online depending on how possible it is to order new servers at OVH. If you are a user of this service, please contact us on the support if you have any questions.
The 2021-03-07 at 19:40 UTC websites on the RBX went down. We started investigating the issue at 19:45 and saw the RBX reverse proxies were not accepting new connections. We restarted them and everything went back to normal by 19:54.
The culprit was a badly configured NOFILE limit on the RBX reverse proxies. We updated the setting accordingly.
Afterwards: We investigated all the reverse proxies on all the zones to make sure the NOFILE limit was correctly configured everywhere. We updated the reverse proxy software (sozu) to refuse to start when given too few NOFILE. We updated the sozu package to enforce the right NOFILE value upon installation.
We experienced an unexpected issue with a core component of the Metrics system.
The service is completely unavailable at the moment. We are working on it.
08:50 UTC: The faulty component is working. We are working on bringing everything back up.
08:59 UTC: Everything is back up. The ingestion pipeline is catching up.
09:07 UTC: The incident is over.
We are investigating issues with our core API.
EDIT 21:07 - fixed.
Some FS-Bucket add-ons will need to be migrated to a different server for security reasons. During this migration, the Buckets will be in Read-Only mode. Any attempt to create or update a file on the add-on will fail, including for FTP operations. Errors related to Read-only file system are expected during this migration.
The migration is expected to last at most 1 hour. All impacted applications will be redeployed during the migration. After the deployment, application will be able to write to the bucket. Read operations will not be impacted.
EDIT: This maintenance has been postponed to 15:00 UTC+1
EDIT 15:00 UTC+1: The maintenance is starting
EDIT 15:02 UTC+1: The buckets are now read-only
EDIT 15:14 UTC+1: Starting now, you can redeploy your applications if you want to regain write access early. Otherwise, affected applications will be redeployed automatically in the upcoming hour, starting with applications of Clever Cloud Premium customers
EDIT 17:14 UTC+1: The deployment queue finished one hour ago, everything has been working fine so far. This maintenance is over
February 2021
On 2021-02-22 at 11:00 UTC, the API and deployment system will go down for a quick maintenance. Expected downtime is up to 10 minutes.
11:00 UTC: Maintenance is starting. Deployments are disabled.
11:02 UTC: API is down.
11:11 UTC: API and deployments are up again. Maintenance is over.
Some Fs-Bucket add-ons will need to be migrated to a different server for security reasons. During this migration, the Buckets will be in Read-Only mode. Any attempt to create or update a file on the add-on will fail, including for FTP operations. Errors related to Read-only file system are expected during this migration.
The migration is expected to last at most 1 hour. All impacted applications will be redeployed during the migration. After the deployment, application will be able to write to the bucket. Read operations will not be impacted.
Emails will be sent to customers of the impacted add-ons.
EDIT 12:00 UTC+1: The maintenance will begin shortly
EDIT 12:04 UTC+1: The buckets are now read only
EDIT 12:13 UTC+1: The redeployment queue began, it should not last more than 15 minutes.
EDIT 12:51 UTC+1: The maintenance is over, the queue ended 20 minutes ago and everything seems to be normal.
We are investigating performance issues with our API.
10:44 UTC: We have found the cause and fixed the issue. It was due to an internal tool unexpectedly making too many costly requests.
We are investigating an issue with Metrics ingestion. Recent data is unavailable at this time.
EDIT 13:18 UTC: Ingestion is working again, working at full speed to catch up.
EDIT 14:03 UTC: Ingestion has caught up since a few minutes ago, everything should be back to normal.
Some Fs-Bucket add-ons will need to be migrated to a different server for security reasons. During this migration, the Buckets will be in Read-Only mode. Any attempt to create or update a file on the add-on will fail, including for FTP operations. Errors related to Read-only file system are expected during this migration.
The migration is expected to last at most 1 hour. All impacted applications will be redeployed during the migration. After the deployment, application will be able to write to the bucket. Read operations will not be impacted.
Emails will be sent to customers of the impacted add-ons.
EDIT 11:55 UTC+1: The maintenance will start on time.
EDIT 12:00 UTC+1: The maintenance is starting
EDIT 12:07 UTC+1: Applications are being restarted. The restart queue should be done in about 20 minutes
EDIT 12:41 UTC+1: The migration is over.
We are experiencing an issue with the storage of application logs. Ingestion is down and read access is partially unavailable.
13:06 UTC: The issue has been solved, ingestion is catching up.
13:10 UTC: Ingestion is all caught up. This incident is over.
Our main API is unresponsive, therefore the console and CLI are unusable as well. We are investigating.
13:53 UTC: The issue is fixed. Everything is back to normal.
An add-on reverse proxy has restarted at 14:04 UTC, leading to connections loss on some add-ons if you used that proxy. Impacted application might have been able to connect to the add-on through a different reverse proxy, unreachable applications will be redeployed.
January 2021
One of the IP (149.56.147.232) of domain.mtl.clever-cloud.com is unreachable because OVH blocked it. We are working on restoring it.
EDIT 18:07 UTC: The IP has been restored. OVH blocked it after a 4 hours email notice of phishing which has escaped our own filters. Further investigations will be conducted to avoid this incident in the future.
Logs are currently delayed and may not be up-to-date. The queue is being consumed. Some messages may have been lost during because of an unexpected service reboot. Logs queries are still working.
EDIT 13:26 UTC: The queue has been consumed. Logs should now be up-to-date.
A FS Buckets server has crashed and failed to automatically restart. An issue was preventing it from properly restarting. It is now fixed.
This server has been unavailable for 8 minutes.
Our shared postgresql leader is currently crashing repeatedly and entering recovery mode, we are investigating what is causing this issue.
Dedicated addons are NOT impacted.