Incident History
Full history of incidents.
February 2019
The main API is having issues with its databases connections, it only ever replies with a 500 error to most requests.
EDIT 07:11 UTC: The issue is resolved
Our deployment system will be unavailable from 11:00 UTC up to 13:00 UTC on Wednesday, 6th of February.
All deployments actions will be queued and started once the deployment stack is back up. The maintenance shouldn't last longer than 2 hours.
Feel free to ask any question on our support regarding this maintenance.
EDIT 11:03 UTC: the maintenance will start soon. Deployments will be shutdown in a few minutes. Push actions on our GIT repositories are disabled.
EDIT 11:06 UTC: Deployments are shutdown
EDIT 11:20 UTC: Deployments should be back, we are still cleaning up things
EDIT 12:20 UTC: We have been keeping a close eye on deployments, everything is going smoothly. Maintenance is over.
Deployments on the Paris zone are currently unavailable. We are investigating the issue and working on bringing them back
UPDATE 9:40 UTC: deployments are back since 20 minutes, we are still cleaning things up.
UPDATE 10:30 UTC: Everything is back to normal, sorry for the issue.
January 2019
We have problems on our deployments systems. We are investigating.
EDIT 8:21 UTC: fixed.
Deployments are currently slowed down. We are working to bring back them to their regular speed
EDIT 18:10: Deployments should be back to normal.
We are running maintenance on live logs and logs drains. There will be some unavailability of these services.
EDIT 11:05 UTC: maintenance is finished.
We will disable MongoDB addon creation and deletion while we are doing a maintenance to add new features. We will edit this post when the maintenance will be finished.
EDIT 16:29 UTC: the new addon dashboard is available. We are continuing the maintenance.
EDIT 17:30 UTC: maintenance finished.
We are experiencing issues on our API which is impacting the console. We are investigating.
EDIT 18/01/19 00:53 UTC: Root issue is most probably identified. The issue was coming from an internal tool. We will investigate this further. In the meantime, the tool has been deactivated and shouldn't cause any harm.
MongoDB free shared cluster is having troubles accepting connections. We are investigating the issue.
EDIT 16/01/2019 09:45 UTC: The problem might be due to old clients drivers being used on the cluster. We have set up a new cluster (version 4.0.3) which should greatly improve things. You can create a new add-on to migrate your database.
To dump your data from your existing, you can use this command: mongodump -u "${MONGODB_ADDON_USER}" -p "${MONGODB_ADDON_PASSWORD}" -h "${MONGODB_ADDON_HOST}" -d "${MONGODB_ADDON_DB}" --archive --gzip
You can then import the data into the new database by using the mongorestore command displayed in the dashboard of your new add-on.
An automatic migration tool for mongodb should be available in the next few days.
One node of the shared RabbitMQ cluster lost its connectivity during 1 minute. It then re-joined the cluster as expected. Real time logs were unavailable at the same time because of that issue.
Redis add-on creation will be disabled starting 13:00 UTC. Dashboard might be unavailable too. This should not include redsmin integration which should remain available.
16:15 UTC: The maintenance is over. Add-on creation and dashboard are now fully available again.
(Hours in UTC)
At 22:27, one of our hypervisors lost access to parts of its disks. Amongst others, It impacted a deprecated front reverse proxy for applications and a front reverse proxy for add-ons (databases). We moved the IP of one of the proxies. The other one, related to the application reverse proxy (62.210.92.244) couldn't be moved and is now unreachable. If you still use it, you should update your DNS records: https://www.clever-cloud.com/doc/admin-console/custom-domain-names/#personal-domain-names
The situation is stabilized. We still consider the infrastructure not fully recovered.
We are adding new features on MySQL Addon. The addon dashboard and management (creation, deletion) will be offline during the maintenance.
EDIT 15:00 UTC: new addon dashboard is available, but addon creation is still unavailable.
EDIT 17.28 UTC: maintenance is now finished.
16:36:30 UTC: A load balancer stops accepting new connections
16:38:00 UTC: An alert due to an important change in network traffic is triggered
16:39:30 UTC: The load balancer is restarted
Everything is back to normal now.
A reverse proxy is dropping some of the TLS connections it receives
EDIT 10:07 UTC: The reverse proxy has been restarted and the issue seems to be resolved. We are monitoring the situation.
The front "mongos" component of the free shared cluster is behaving erratically. We are investigating it.
EDIT: There was sudden drops in free disk space. We change the logging method and it seems to have stabilized the system. We are still working on figuring out the issue.
A maintenance operation is in progress on the Europe MongoDB shared cluster.
We are having issues with the authentication component. Open connections are working fine, new connections are impossible for now.
17:21 UTC: It should be fixed. We are making sure.
17:30 UTC: Incident over.
December 2018
Two reverse proxies are having intermittent networking failures. Those reverse proxies are only when your domain is configured to use A records. Domains using CNAME records should be reachable as usual. We are working on it
EDIT 20:15 UTC: Incident resolved, it was due to a network miss-configuration. We will ensure this doesn't reproduce anymore.
Deployments will be unavailable for up to 30 minutes starting 13:00 UTC because of a maintenance on our deployment system. Deployment actions like START, RESTART, STOP, ... will be unavailable but will remain in queue and will be processed at the end of the maintenance.
EDIT 13:06 UTC: The maintenance is starting EDIT 13:17 UTC: Deployments are now available again. Queued deployments have been processed.
Maintenance is over.
The Clever Cloud API is currently down, we are investigating.
EDIT 16:53 UTC: API is fixed. We detected a problem on our reverse proxies, we are currently fixing it.
EDIT 16:54 UTC: fixed.