Incident History
Full history of incidents.
April 2019
We are experiencing issues on applications/addons access. We are now investigating and we will come back with more informations.
EDIT 17:56 UTC: the systems are backing to normal. It was a DNS resolver problem.
EDIT 17:58 UTC: fixed.
There is a problem that prevents addon migration (if you start or have a running migration, the process will fallback to previous state without problems).
The cleverapps.io domain has been marked as dangerous. As far as we know, one or more subdomains have been reported as dangerous and the complete domain has been added to the list.
This means that your browser may show you a security alert when visiting a cleverapps.io site.
We are looking into reporting the mistake to the relevant lists and services.
Meanwhile, we remind our users that they should never use a cleverapps.io domain for production; they should only be used for development and tests.
March 2019
A cellar-c1 node crashed in a way which caused a very important load on all other nodes; this is causing general slowness and an elevated error rate.
It should go back to normal gradually and will not take more than an hour at the most.
EDIT 17:00 UTC: Error rate and performance is back to normal
We are experiencing issues on our Cellar features.
EDIT 7:25 UTC: multipart uploads are down, the fixes are ongoing. EDIT 15:38 UTC: the cluster has been fixed, everything is back to normal.
We are experiencing a network issue on the older infrastructure in Paris. We are investigating.
EDIT 06:12 UTC: The network issue is over. This was an issue with our provider which affected all our servers but not all at the same time. Nothing was actually fully unreachable at any point in time but there was a lot of packet loss.
We are getting reports from some SFR network users who cannot access the Clever Cloud Console. It seems to impact only some SFR customers.
EDIT 9:19 UTC: This only affects the older SFR network, not the SFR-Numericable network. This specifically affects all SFR peering going through TH2.
EDIT 9:50 UTC: This has been resolved at 9:36:30; if you are still experiencing issues, please tell us.
A network issue is happening. Applications may be unreachable.
Console is partly down. Some apis are down.
EDIT 18:20 UTC: Here is the history and context of the network issue:
At 17:25, a maintenance on a component of a redundant network link caused one of the underlying links to fail. For reasons unknown at this time, the failing link was elected and about 30% of packets were lost until 17:29.
At 17:30, the network engineer decided to revert the change; this caused additional loss for about 30 seconds. Network was back to normal at 17:31.
We are investigating an elevated error rate and elevated response times on Cellar. Only some buckets / files are affected by this issue.
EDIT 14:01 UTC: Error rate is back to normal. Response times are going down, we are still watching the situation closely.
EDIT 15:40 UTC: We are seeing an elevated error rate again, this was caused by a restart of a node which triggered a very high load on other nodes (which is not supposed to happen). We are investigating.
EDIT 16:30 UTC: The error rate went down significantly but it's not over yet. We sadly cannot give any meaningful ETA as of now.
EDIT 16:55 UTC: The error rate is close to normal. One node is still in trouble and it's causing a few errors; it should resolve quickly.
EDIT 17:15 UTC: The failing node went back to normal at 17:02. We are still seeing a few errors for write requests as of now.
EDIT 17:23 UTC: The error rate is back to normal. A few nodes are still a bit slower than usual so performance is a bit hit or miss but it should go completely back to normal in up to an hour.
February 2019
We are experiencing TLS issues on some HTTPS requests.
EDIT 15:33UTC: fixed.
Intermittent network issues have been identified affecting several systems. Those issues resulted in various timeouts or longer than expected connections to databases or applications.
We didn't see any new timeout since 23:45 UTC but we continue to monitor the service.
Live logs are currently unavailable. Newer logs should be available by refreshing the logs panel. Logs drains may be impacted to.
EDIT 10:30UTC: fixed.
Due to a incident one Online Datacenter: https://status.scaleway.com/incident/286, we are experiencing issues.
EDIT 21:00 UTC: fixed.
An Hypervisor restarted. Applications have been redeployed and add-ons are restarted.
EDIT 15:28 UTC: All add-ons should be back online, some of them took longer than expected to recover. The cause of the reboot will be investigated.
Deployments are currently unavailable due to an ongoing issue with our deployment system
EDIT 16:42 UTC: The root cause has been found. We are redeploying core components to clean everything.
EDIT 16:50 UTC: Deployments are available since a few minutes now. We are still cleaning things up. Sorry about the issue
We are experiencing issues on one hypervisors which can impact addons.
EDIT 20:50UTC: we are hard rebooting the hypervisor.
EDIT 20:55UTC: the hypervisor is up, the addons hosted on it are starting.
EDIT 20:58UTC: fixed.
Some deployments are having a delay to start.
10:27 UTC: The issue is now fixed
Applications in the console are reporting an "unknown" state.
We are investigating the issue. Deployments are stopped until we find the root cause.
EDIT 16:35 UTC: Applications state should now be OK. Deployments are still stopped until we figure out the issue.
EDIT 16:40 UTC: Problem has been identified. We will resume deployments in a few minutes. All deployments action were queued and will be consumed.
EDIT 16:42 UTC: Deployments are enabled again. It may take a few minutes before your actions are handled. We consider this incident over.
We have an abnormal amount of 503 errors served by our reverse proxies. We are investigating.
EDIT 11:31 UTC: Cause has been identified, we are currently fixing the issue on our reverse proxies.
EDIT 11:34 UTC: All reverse proxies now have a consistent state. The issue is fixed.
The issue happened after a configuration error made during a manual operation on some of the reverse proxies. Applications that redeployed since 11:08 UTC were impacted by that issue. Other applications were fine. The changes were rollbacked and will again be tested thoroughly on our test infrastructure.
A maintenance is scheduled on Monday 2019-02-18 at 11:00 UTC (12:00 noon, Paris time (CET)); it will affect the main API, deployments and GIT repositories.
The maintenance will last at least 5 minutes but no more than 20 minutes.
Different parts of the system will be affected throughout this maintenance, please wait until the end of the maintenance before reporting any issues you may be having.
EDIT 11:00 UTC: The maintenance will start in a few minutes. Deployments and GIT repositories will be unavailable. The console might report an "unknown" or not up-to-date state for applications. This is expected.
EDIT 11:05 UTC: Maintenance is starting, deployments are down and so are GIT repositories (push actions will be rejected)
EDIT 11:09 UTC: Deployments are available again. Push actions on GIT repositories are still disabled.
EDIT 11:10 UTC: Our main API is entering read-only mode. 500 errors might appear during this time.
EDIT 11:11 UTC: Git repositories are now available. You might need to clear your DNS cache to be able to push again.
EDIT 11:20 UTC: Our main API should be fully available again. We are looking if everything looks fine.
EDIT 11:23 UTC: Everything is looking fine. The maintenance is over. You might experience git push errors up until 45 minutes. To avoid that, please clear your DNS cache.