Incident History
Full history of incidents.
October 2018
Deployments actions (start, restart, stop, git push, ...) are slower than usual. We are looking into this.
13:09 UTC: We are going to restart one of the deployment core system. Deployments actions (like the one above) will be unavailable for up to 30 minutes. All actions will be queued and executed at the end of the maintenance.
13:40 UTC: Another problem occurred during the restart of that system. We are now trying to fix this one.
EDIT 14:03 UTC: Deployments are available since ~5 minutes now. We are still cleaning things up before closing this incident.
EDIT 14:30 UTC: Everything should be back to normal now. Sorry for the extra maintenance time and the deployments unavailability.
Our monitoring system had a network cut making it see a lot of applications unreachable. Those applications are being redeployed but it may add delays for new deployments (start / redeploy / stop) because of the number of ongoing deployments.
EDIT 12:50 UTC: The deployments should now be back to normal. Apologies for the delays.
September 2018
The GitHub API has changed so we are patching our API to fix the auto-deployments of new applications, the fix will be retroactive.
The follower stopped replicating data and taking backups.
We are trying to restart it.
Postgresql leader is down. Promoting follower and update domains.
DNS has been updated. Clients should connect back to the database
EDIT 12:22 UTC: The new leader is correctly serving requests since 0:30 AM UTC.
The shared cluster is experiencing performance issues. We are working to mitigate those issues.
Deployments using cache (build cache, dependencies cache) are failing because the cache can't be downloaded. We are investigating
EDIT 10:17 UTC: We are still working on the issue. If you have troubles deploying, you can set your application's scalability settings to which a dedicated build instance would use. Do not hesitate to ping our support if needed.
EDIT 10:25 UTC: ETA is 2 hours if everything goes well.
EDIT 12:30 UTC: The deployments with cache are back. Everything should work as expected from now. Sorry for any failed deployments or longer than expected deployment times.
One of the two reverse proxy of *.cleverapps.io crashed and had a longer than usual restart time. Traffic hitting this server didn't complete as requests would hang until the connection timed out.
The problem has been resolved at 16:08 UTC
One of the nodes of the shared rabbitmq cluster crashed. It's currently restarting.
EDIT 18:50 UTC: The node has successfully restarted, the cluster should now be operational as usual
August 2018
The main API is unavailable, the console cannot be loaded as well.
We are looking into it.
EDIT 15:30 UTC: Our API is back online. The console can now be loaded.
An hypervisor is unreachable, we are working on fixing the issue.
Applications on this hypervisor are being automatically redeployed. Add-ons are unreachable.
EDIT 12:21 UTC: The hypervisor is back online and is restarting the add-ons.
EDIT 12:32 UTC: All add-ons are now reachable.
Deployments will be interrupted during 30 minutes at 12:30 UTC+2 today. A core component upgrade will be performed. This will not impact already running applications or add-ons. All deployments will be queued and executed at the end of the maintenance.
The maintenance shouldn't last longer than 30 minutes but it may be possible that some delays occur. We will update this ticket to let you know about the status of the maintenance.
EDIT 12:25 UTC+2: New deployments are stopped to be consumed.
EDIT 12:30 UTC+2: The maintenance has started
EDIT 12:56 UTC+2: Deployments are back since ~10 minutes. We are still cleaning things up
EDIT 13:03 UTC+2: Maintenance is over and was successful. Do not hesitate to contact us if anything's wrong on your side.
Deployments are temporarily disabled as we fixed the issue with a component of the deployment system.
EDIT 19:17 UTC: This was actually a false positive from our monitoring. After verifying that the component is working fine and fixing the monitoring probe, we re-enabled deployments.
One FS Buckets server is unavailable, we are awaiting news from our provider.
EDIT 05:28 UTC: The server is partially and randomly available: the problem has been identified by our provider: it's coming from the switch the server is connected to. They are working on fixing the issue.
EDIT 08:04 UTC: Issue is fully fixed since 07:30 UTC
MongoDB cluster will not accept writes until failure is fixed.
The failing node is up again.
Creation of add-ons and buckets on cellar is temporarily failing. We are working on it
EDIT 15:30 UTC: The creation of add-ons and bucket is now fixed. It may take a little longer than usual but these slowness will be resolved in a few hours
There is an issue with the entry point to the cluster.
Users are stretching the "fair usage" concept way above reasonnable limits. We are working with them to enforce the fair usage.
Performance should have been restored.
We are still watching the cluster.
A network issue is preventing the logs system from working.
EDIT 13:17 UTC: Logs should be available, the cluster is slowly recovering
EDIT 13:23 UTC: The logs cluster is UP and running again, logs shouldn't have been lost thanks to buffering.
Sorry about the inconvenience.
A maintenance of our Git repositories will be held on Thursday (2018-08-09) at 1pm, UTC + 2.
Write operations like "git push" or "clever deploy" to Clever Cloud repositories won't be possible during 30min. Read access won't be affected during this time.
Thanks for your patience.
EDIT 13:00 UTC+2: The maintenance is starting
EDIT 13:05 UTC+2: The maintenance is now complete. Do not hesitate to open a support ticket if anything goes wrong. Thanks for your patience!