Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

December 2022

Fixed · Infrastructure · Global

Following a maintenance, few servers that host some dedicated load balancers are seen unreachable by the monitoring

EDIT: 2022-12-20 19:47 UTC : During the recovery process some services goes down with tls issues

Paris: an HV is down
Fixed · Infrastructure · Global

An hypervisor in Paris is not responding. We are investigating.

Update 4:16 AM UTC: HV is now up. We are running the cleanup tasks associated to the HVs

Update: 4:54 AM UTC: Cleanup is over.

Paris: networking issues
Fixed · Infrastructure · Global

One of our Paris datacenter encounters network issues. Some servers were unreachable for about 1 minute at 10:09 UTC. Services (applications, add-ons, Cellar, ...) hosted on those servers were partially unreachable or fully unreachable during that time (depending on the scalability or replication of those services).

The cause has been identified and a solution is currently being investigated. This incident will be updated as soon as we have more information.

EDIT 13:44 UTC: Another network interruption happened at 13:01 UTC. A fix is currently being tested.

EDIT 14 Dec 2022 15:55 UTC: The fix appears to be working as expected. This incident is now over.

Fixed · Infrastructure · Global

Some applications are having troubles being monitored by our systems and get redeployed with the Monitoring/Unreachable reason even though they are still available. We are investigating the cause of the issue.

EDIT 15:55 UTC: The cause has been found. This issue only affects applications tied to a unique IP proxy service. The issue has been mitigated in the last minutes and we are working to fully fix it.

EDIT 16:20 UTC: The issue has been fixed and should not happen again. If you encounter weird Monitoring/Unreachable deployments, feel free to contact our support team.

Deployment issues
Fixed · Deployments · Global

Our distributed deployment message queuing system lost a broker which has lead to some lags during the consumption and a few components unable to reconnect properly to the cluster.

EDIT: 18:30 UTC - all systems are up

November 2022

Fixed · Cellar · Global

We are experiencing high read-write latency on Cellar. We are working on it.

EDIT 24 of november 9:33 UTC: Balancing is over.

Deployment lag
Fixed · Deployments · Global

Some deployments lag

A monitoring desynchronization causes disturbance on deployments. We are investigating the troubles. We are manually cleaning unnecessary deployments

We have to clean some stuck deployment but system is now recovered

Pulsar production issue
Fixed · Pulsar · Global

The storage layer does not allow any writes more. We are scaling the Pulsar storage capacity

Fixed · API · Global

We are currently having an increased response time on our main API (api.clever-cloud.com). Services relying on it are impacted and might experience extra time to answer to requests.

We are investigating the issue.

EDIT 14:26 UTC: The underlying issue has been identified and fixed. Services, including the Console and CLI should now be loading as usual. Sorry for the inconvenience.

Fixed · Access Logs · Global

We are currently having ingestion issues on the metrics and access logs services. Data for the last 20 minutes is currently missing. We are investigating.

EDIT 11:23 UTC: Ingestion lag is now resolved, metrics and access logs should now be up-to-date.

October 2022

SYD/SGP network issues
Fixed · Infrastructure · Global

Our monitoring has raised network issues, we are investigating.

Status:

  • Ping does not go through between PAR (on BSO Network) and SGP/SYD (OVH Network)
  • Ping does go through between PAR and other OVH zones (RBX, MTL, WSW…)
  • Ping does go through between RBX and SGP/SYD (OVH to OVH)
  • Applications on both SGP and SYD are still UP and reachable from other networks. Deployments on these zones are still unavailable.

UPDATE 20:13 UTC: Network is kind of coming back up, but we see 80% to 90% packet loss. 21:50 UTC: still a lot (90%) of loss on the PAR -> SGP/SYD route, way less (30%) in the SGP/SYD -> PAR route. 2022-11-01 0812 UTC: >90% of loss on the PAR -> SGP/SYD route. 2022-11-01 1812 UTC: network seems fine.

Fixed · Reverse Proxies · Global

PAR reverse proxies unavailable for a short time period

Network issue on OVH
Fixed · Reverse Proxies · Global

Some network issues on our provider OVH can leads to some desync of reverse proxy configurations

Fixed · Access Logs · Global

We are experiencing performance issues on our metrics/accesslogs infrastructure. We are on it.

Update 10:36 UTC: Performance has been fixed.

Fixed · Infrastructure · Global

Some deployments can have abnormal delays to deploy. Applications may experience slowness.

** EDIT 11:55 UTC **: We have found the root cause, we have mitigated the issue. we are deploying the solution.

Fixed · Access Logs · Global

The data store behind Metrics and access logs have lost a node. Some lags to query metrics and access logs could be observed.

Issues on OVH network
Fixed · Infrastructure · Global

At 11:26 UTC, the monitoring started alerting about unreachable proxies on RBX, RBX-HDS, MTL2, WSW. All these zones are hosted on the OVH network.

We are investigating and watching the situation.

At 11:53 UTC, the monitoring sees everything up again. We perform a few check on some services

Deployments slowness issue
Fixed · Deployments · Global

We observe slow deployment times, we are investigating why.

** EDIT 18:10 UTC ** : The issue has been identified and actions to solve this issue has been performed

Deployments issues
Fixed · Deployments · Global

Due to the pulsar incident, some deployments may fail from time to time.

Some hypervisors are behaving strangely. We are watching and fixing them.

EDIT 10:20:00 UTC: Deployments are currently unavailable while we work around the issue.

EDIT 11:31:00 UTC: Deployments issues are fixed. We continue to monitor the situation. If you have troubles redeploying an application, please contact our support.

POSTMORTEM: The Pulsar outage that started around 04:30 UTC (see https://www.clevercloudstatus.com/incident/574) got in the way of:

  • the deployment process, breaking some notifications at 09:30 UTC.
  • the uptime of some persistent VMs (like databases) (See https://www.clevercloudstatus.com/incident/576), making the monitoring trigger deployments.

The pulsar notification system is being gradually deployed on our infrastructure, having passed the tests on our preproduction zone. We do have a fallback method for notifications. However, the issue was weird enough that the pulsar notification was not cleanly failing. They rather timed out after a long time, preventing the fallback to trigger. We stopped all deployments at 10:20 UTC. We worked on quickly adding an emergency flag to prevent the hypervisors from using pulsar for notifications. This way, we can bypass it and go straight to the fallback method.

To avoid this issue, we are working on the following:

  • monitor the pulsar logs before it impacts the rest of the production.
  • try to mitigate the long timeout issue on the notification actors, allowing for a quicker fallback.
Pulsar add-ons issues
Fixed · Pulsar · Global

The pulsar cluster hosting the pulsar add-ons is undergoing issues. We are investigating.

POSTMORTEM (all times are UTC) : Around 04:30: Timeouts in inter-nodes connections started to show up in the logs. They did not lead to alerts in the monitoring Around 05:00: We start getting issues in our infrastructure from software using that cluster.

11:30 : we disable the brokers to analyze the issue.

14:42: The incident is now resolved. If you still encounter any problems, please contact our support.