Incident History
Full history of incidents.
January 2018
Network instability on Online DC2 makes some products unreachable:
- Mysql shared cluster
- Postgresql shared cluster
- Mongodb shared cluster
- One of the cleverapps front proxies
The shared mongodb cluster is experiencing issues, we're working on bringing it back up.
Due to disk space, we need to lower the number of logs we store, for now. Only the last 4 days are kept, instead of the ideal number of last 7 days.
EDIT 2018-06-15 UTC: All 7 days are now available again.
December 2017
A core component will be upgraded. Deployments will be disabled for an hour starting at 11:30 UTC. This upgrade should fix some deployments delay among other things.
EDIT 11:31 UTC: Maintenance is starting
EDIT 12:06 UTC: Deployments are back, we are now cleaning some old artefacts
EDIT 13:00 UTC: The maintenance is over
Our deployment system encounter some slow down. Some application may take longer than usual to deploy. We are working on it
EDIT 19:25 UTC: Those slow downs might require an infrastructure change that will be done next week. Until then, slow downs should be less frequent and less important
EDIT 2017-12-08: 12:00 UTC: Deployments take less time after some fixes on our end. The migration will still happen to entirely fix it. Incident is considered as closed because we don't see any more extra times.
We've observed an elevated error rate on two front load balancers newly added to the pool. We're pulling traffic back from these load balancers.
November 2017
Some deployments might have troubles starting a deployment. We are investigating.
EDIT 17h31 UTC: Deployments are disabled for now EDIT 17h38 UTC: Deployments are now back up but may be stopped again in a few minutes if needed EDIT 17h55 UTC: The incident is now resolved. We will keep an eye on it for the upcoming days
We are experiencing a network issue on one of our front. The support team is actively working on this.
EDIT 14:56 UTC+1: Unreachable servers are being restarted and will be available shortly. In the meantimes, impacted applications are being redeployed
EDIT 15:26 UTC+1: The team is performing the final cleanup. The issue is about to be closed. The remaining apps and add-ons are being restarted.
EDIT 15:50 UTC+1: The outage is now resolved. Contact the support is you encounter any trouble.
Due to a software update, deployements will be disabled for up to 30 minutes starting at 12:30 UTC+1
EDIT 13:00 UTC+1: The maintenance is over, deployments are back since 15 minutes
A network issue affecting a front load balancer on the PAR zone has been identified and fixed
October 2017
Due to a network issue, some FS buckets have been unavailable for a short period of time. All FS buckets are now available.
One hypervisor has experienced a hardware issue and is rebooting. Affected apps are being redeployed, affected addons will be available shortly.
Due to a phishing application deployed on a cleverapps.io domain, the whole domain name has been marked as malicious.
We are working on clearing the alert. In the meantime, we'd like to warn you that cleverapps.io domain names are provided only for test purposes and that they should not be used in production.
September 2017
A node hosting shared redis databases has been restarted after having connectivity issues. Impacted applications will automatically be redeployed when connectivity is restored.
Our main API is currently slower than usual. We are looking into it
EDIT 12:09 UTC: The API is now performing smoothly. We will keep looking why it went into such state
A shared redis cluster went down. It's being restarted
EDIT 17:30 UTC: all shared redis are now available again
The .cleverapps.io domains have issues resolving through multiple DNS servers. It seems like the top .io TLD DNS servers are the root cause of the problem. Users using this domain may have error messages like "Server not found".
If you need it, here is the IP of the domain: 217.70.184.38
EDIT 19:43 UTC: The incident seems to be resolved, .cleverapps.io domains now resolve correctly
One of our hypervisors is having a huge load, making it unresponsive.
Impacted applications are being redeployed
EDIT 10:15 UTC: The server is still under huge load. Services on it continue to answer correctly in most cases. Applications are still redeploying
EDIT 10:30 UTC: The server is now reachable and responsive, we are looking into why it went under such a heavy load
One of our physical server has gone down. Impacted applications are being redeployed and we are investigating the incident
EDIT 14:24 UTC: The server is still down, we are waiting for more informations from our prodiver
EDIT 14:37 UTC: One of the server's fan has died and the server won't start.
EDIT 14:43 UTC: Impacted databases will be migrated on another server on request to the support. We will also contact impacted users. Let us know if you want to start a new database using tonight's backup
EDIT 15:23 UTC: Our provider is replacing the fans, no ETA for now
EDIT 16:55 UTC: Our provider replaced the fans and the server is now back up. Non migrated databases have been started again and linked applications are being redeployed. We will continue to monitor the situation
A reverse proxy serving addons traffic has started refusing connections. It has been restarted and is now serving traffic correctly. The affected applications have been automatically restarted.