Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

December 2018

PostgreSQL Addon
Fixed · Console · Global

The PostgreSQL Addon Dashboard is currently unavailable, we are working to fix it.

EDIT 15:17 UTC: fixed.

Fixed · SSH Gateway · Global

The SSH gateway is currently unavailable. We are working on bringing it back as soon as possible

EDIT 12:18 UTC: We are still trying to figure out a fix for the issue.

EDIT 12:47 UTC: The problem should now be fixed. A configuration error made this incident longer than it should have last. Applications may need to be redeployed to get the SSH service back online.

Sorry about this incident.

Unresponsive reverse proxy
Fixed · Infrastructure · Global

One of our reverse proxies went quite unresponsive but was still able to process some requests and report its state to our monitoring. Most of the requests it received weren't processed. This is now fixed.

Sorry for the inconvenience

Fixed · MySQL shared cluster · Global

A MySQL shared cluster is overloaded at the moment. We are looking into which users are over-using it.

14:26 UTC: One culprit has been found. The cluster's load has been reduced significantly.

14:38 UTC: The cluster's load is back to normal since 14:30.

Deployments issue
Fixed · Deployments · Global

Some deployments fail to start and/or are not being properly reported by the API. We are investigating.

08:38 UTC: We are restarting part of the deployment system.

08:49 UTC: Since 5 minutes ago, deployments are being processed with some delay.

08:54 UTC: Back to normal.

API unavailability
Fixed · API · Global

Some API endpoints seem to be currently unavailable making the console unavailable too. We are currently investigating what's causing this.

10:00 UTC: We found the root cause. The console still can't be loaded at the moment but other services should now be available (like deployments) 10:06 UTC: There was an underlying issue causing the console loading. It is now fixed. The incident is now over. Sorry for the inconvenience

Metrics unavailability
Fixed · Access Logs · Global

Metrics are currently unavailable for read requests. Write requests are working as expected.

EDIT 14:04 UTC: Metrics are getting back up

EDIT 14:10 UTC: Metrics are fully recovered. Sorry for the inconvenience

November 2018

Fixed · Deployments · Global

As precised on Cogentco status page,

Cogent will be performing code upgrades in the following areas.
During these upgrades, customers in or transiting the area may experience
intermittent periods of packet loss and latency between 15 and 45 minutes 
for the duration of the window.

Location: Paris, France
Start time: 11/30 00:01 CET
End time: 11/30 06:00 CET
Work order number: NC840-119

our link with Montréal (MTL) zone can be affected by issues, so our systems (deployments, monitoring, etc.) on Montréal (MTL) can experiences issues.

Fixed · Deployments · Global

We have experienced issues on deployments on Montréal (MTL) zone between (from 15:24 to 15:30 UTC).

Issue with webhooks
Fixed · API · Global

A validation test of an update to the webhook API has made its way to production. Clients received events not meant for them for 8 minutes. This is now fixed.

Sorry for the inconvenience.

Upgrade PostgreSQL addon.
Fixed · API · Global

We are deploying a new feature on PostgreSQL addon, the creation and management of those addons is currently disabled.

Fixed · Access Logs · Global

A core component keeps restarting making the metrics unavailable for fetch. No metrics are lost during those restarts. We will take actions to fix this issue in the upcoming days.

An action was taken at 02:30 UTC (2018-11-21) which has successfully fixed this issue. This is only temporary though.

A permanent fix will be applied later today, which will require a downtime of that component.

EDIT 2018-11-21 16:50 UTC: The permanent fix is delayed to tomorrow, 2018-11-22.

EDIT 2018-11-22 10:40 UTC: The fix will be applied at 10:50 UTC, this will require at least one restart of that component which will lead to an unavailabiliy of Metrics for about 20 minutes.

EDIT 2018-11-22 11:25 UTC: Metrics are back since 11:08 UTC. Incident over.

Network issue
Fixed · Infrastructure · Global

A network issue (apparently) is affecting several hypervisors and services. We are investigating.

EDIT 19:21 UTC: Here is the incident of our provider: https://status.online.net/incident/153 (3 racks have lost public connectivity)

EDIT 20:33 UTC: The issue should be fixed. As of now, our monitoring is happy. We are cleaning up.

Fixed · Infrastructure · Global

It has been reported that some database are slower than usual because of network slowness. We investigated and took actions against our reverse proxies. One of them has been fully restarted leading to loss of established connections. We are currently monitoring if those actions are improving the situation.

EDIT 12:10 UTC: The issue seems to be resolved now

Fixed · Access Logs · Global

Metrics currently can't be accessed. Metrics ingestion still works, only metrics fetching will not work.

EDIT 16:25 UTC: One of the component was failing due to a network configuration error. The network configuration has been fixed and the component is currently restarting. It should be restarted in about 15 minutes.

EDIT 16:40 UTC: The component has restarted, metrics are now available again for read actions. No data was lost. Sorry for the extended interruption.

Fixed · Infrastructure · Global

A network issue is currently happening on our reverse proxies on the Montreal zone. We are currently working on it.

EDIT 13:28 UTC: The network issue has been resolved since 13:20 UTC. Everything should be back to normal. Sorry for those issues.

Fixed · API · Global

Our API was unavailable for 10 minutes. The CleverCloud console couldn't load and deployments wouldn't start. This has been fixed.

October 2018

Hypervisor unreachable
Fixed · Infrastructure · Global

A hypervisor is unreachable.

Affected applications are being restarted automatically.

Affected addons are unreachable.

EDIT 17:56 UTC: Looks like it's a network issue, we are awaiting word from our provider.

EDIT 18:08 UTC: Our provider tells us they are working on it, no ETA nor details given.

EDIT 18:26 UTC: There was a short electrical outage in the datacenter where this server is, some routers and switches have been impacted by the switch to the backup power source. They are working on fixing affected network hardware.

EDIT 18:44 UTC: The server is back, addons should be reachable. We are making sure that everything is back online.

EDIT 18:56 UTC: Everything is working fine. Incident closed.

Fixed · RabbitMQ shared cluster · Global

Some of the nodes of the cluster crashed. They are currently being restarted. Users using this cluster may experience disconnections and failures to read / publish messages.

Update 16:34 UTC: The cluster nodes have been restarted. The cluster is UP again. Sorry for the inconvenience.

Fixed · Infrastructure · Global

A human error caused an issue with the configuration of the add-ons reverse proxies at 12:18 UTC. MySQL dedicated add-ons were unavailable at this point, except for already open connections.

At 12:30 UTC, we found the cause of the issue.

At 12:32 UTC, the issue was fixed and we regenerated the reverse proxies configuration.

At 12:33 UTC, add-ons were available again.

We have put the necessary protections in place to prevent this from happening in the future.