Skip to main content
Clever Cloud Status

Incident History

Full history of incidents.

Newest first

October 2024

Fixed · cleverapps.io domains · Global

The certificate of cleverapps.io has not been properly renewed at 12h06 UTC. A manual regeneration of the certificate in on the way.

EDIT: The certificate has been renewed at 12h33 UTC, it has been applied and propagated to all load-balancers.

Fixed · Global

A scheduled network maintenance will be carried out in the Paris region on Wednesday, October 2, 2024. This upgrade will affect non-production links, and no impact on production systems is expected.

Start Date & Time: 2024-10-02 20:00 UTC

End Date & Time: 2024-10-02 21:00 UTC

We will provide regular updates throughout the maintenance period.

EDIT 20:38 UTC: The maintenance is now starting.

EDIT 21:45 UTC: The maintenance is still ongoing. Most of the operations are over, verification are currently taking place.

EDIT 22:20 UTC: The maintenance is now over. No impact detected.

Fixed · Deployments · Global

We are experiencing issues with the deployment pipeline.

EDIT 12:43 UTC: the system has returned to normal operation. Our team is continuing to investigate the root cause to ensure stability moving forward. Further updates will be provided as necessary.

EDIT 13:12 UTC: fixed.

September 2024

Pulsar instabilities
Fixed · Pulsar · Global

Following a maintenance operation to reduce load on pulsar cluster. The cluster has an the issue with some configurations, we are investigating the reason.

Platform instability
Fixed · Infrastructure · Global

We have detected some latency and a few instabilities to connect to our platform, we are investigating.

EDIT 17h40 - Root cause has been identified, and network is now stabilized. We are closely monitoring the platform to be sure this incident is closed

Fixed · Services Logs · Global

We detected an issue on log reads.

EDIT 13:00 UTC: identified and patched. We are currently deploying the fix.

EDIT 13:15 UTC: fixed.

Pulsar instability
Fixed · Pulsar · Global

We are experiencing a pulsar outage, which impacts logs and access logs and other components of the platform. Preliminary root cause seems like a zookeeper problem. We are working on it.

EDIT Fri Sep 20 18:16:00 2024 UTC Deployments have been disabled. We are still investigating the Zookeeper outage, causing Pulsar outage.

EDIT Fri Sep 20 19:57:09 2024 UTC: The zookeeper quorum is back online, and therefore Pulsar. Deployments have been enabled, we are watching the situation.

EDIT Fri Sep 20 22:09:40 2024 UTC: Pulsar cluster is still unstable, deployment have been disabled.

EDIT Fri Sep 20 23:36:10 2024 UTC: Deployments queue is back, we are ramping up logs's data usage to avoid bursting Pulsar too much.

EDIT Sat Sep 21 00:49:10 2024 UTC: Pulsar cluster is now stable. Applications should now have their logs available in the console / CLI as well as the drains. Access logs lag is currently catching up. We continue to monitor the situation.

Pulsar instability
Fixed · Pulsar · Global

We are experiencing a pulsar outage, which impacts logs and access logs and other components of the platform. Preliminary root cause seems like a zookeeper problem. We are working on it.

EDIT Thu Sep 19 20:49:09 2024 UTC: since 20:20, ZK quorum is up, and all services connected to Pulsar are now back online

EDIT Fri Sep 20 07:43:00 2024 UTC: we are still impacting by zookeeper outage, we are investigating the issue, the logs and access logs stack are currently unavailable

EDIT Fri Sep 20 08:04:00 2024 UTC: we have found the issue on pulsar side that was trying to write indefinitely metadata on zookeeper. we have restarted the broker that had the issue. We are watching, the situation is going back to normal

EDIT Fri Sep 20 08:20:00 2024 UTC: we are still watching the metrics from the pulsar cluster, the situation is going back to normal. we are recoverying from lag on the access logs ingestion, current eta is around 12:30 utc.

EDIT Fri Sep 20 13:15:00 2024 UTC: we have fully ingested the access logs, the cluster pulsar is working normally.

[WSW] region instability
Fixed · Infrastructure · Global

Wed Sep 18 22:22:29 2024 UTC: Several hypervisors have been rebooted in WSW. They came back 40min ago, and we are fixing several services who are not online.

EDIT Wed Sep 18 22:30:47 2024 UTC: we have been impacted by https://bare-metal-servers.status-ovhcloud.com/incidents/j7f4kpv9f17z. All services are now online

[Paris] Network upgrade
Fixed · Global

On September 18, 2024, our network provider will carry operations to improve network resiliency on the Paris region. No service interruption is to be expected during that upgrade. This is a follow up of https://www.clevercloudstatus.com/incident/893.

Start date: 2024-09-18 19:00 UTC

End date: 2024-09-18 23:00 UTC

EDIT 2024-09-18 19:18 UTC: The maintenance is starting.

EDIT 2024-09-18 20:53 UTC: The maintenance is now over. No service interruptions noted.

Fixed · Infrastructure · Global

At 06:41 UTC, we got an alert that all the WSW region stopped responding. At 06:44 UTC, we got hold on the hypervisors. The first check showed they had been rebooted. At 06:50 UTC, all customers services were up and running. At 07:15 UTC, we finished all the checks that the region is fine.

Here’s the matching OVHCloud status: https://bare-metal-servers.status-ovhcloud.com/incidents/hw285l60sq7h It looks like an electrical incident happened on the racks that hold our servers.

Fixed · Infrastructure · Global

Our monitoring system has report us high latencies to interact with SYD and SGP region. We are investigating the issue.

EDIT 08:50 UTC : The latencies goes back to normal, we are still watching the issue.

[Paris] Network Issues
Fixed · Infrastructure · Global

We are experiencing network issues on the Paris region and are working to identify them.

EDIT 18:21 UTC: the situation seems back to normal. We are still working to identify the reasons;

EDIT 18:23 UTC: we are working to restore impacted components.

EDIT 18:28 UTC: while preparing an intervention in one of our data centers in Paris, we encountered an unfortunate network rerouting. Services are now fully operational again.

EDIT 20:40 UTC: Updated wording to include "Paris region" for impacted location.

Fixed · Deployments · Global

The build cache upload of deployments has an elevated error rate since 19:05 UTC. The root cause has been identified. This may prevent your deployments to finish correctly.

EDIT 22:25 UTC: The service is now fully operational again. Builds that failed because of this issue should be restarted. Please contact our support team if you need any assistance.

[Paris] Network upgrade
Fixed · Global

On September 11, 2024, our network provider will carry operations to improve network resiliency on the Paris region. No service interruption is to be expected during that upgrade.

Start date: 2024-09-11 19:00 UTC

End date: 2024-09-11 23:00 UTC

EDIT 2024-09-11 19:36 UTC: The maintenance is starting.

EDIT 2024-09-11 23:00 UTC: The maintenance is now over. No additional impact besides the ones described in the following incident: https://www.clevercloudstatus.com/incident/895

Fixed · Infrastructure · Global

At 17:00 UTC, a hypervisor (hv-mtl2-012) stopped responding. The on-call team got an alert and starting the investigation. It seems that the hypervisor just rebooted itself.

We are trying to find the reason and making sure that all the services on that server restarted correctly.

UPDATE 17:24 UTC: the team just finished checking all the services: they are now up and running.

update: OVHCloud’s status confirms what we saw (server rebooting for no reason): The problem impacts other servers (not ours) as well. Fortunately for us, we made sure to avoid choosing our OVH servers in the same racks. We’ll wait for the result of their investigation.

UPDATE 2024-09-05 08:55 UTC: The incident has been resolved on OVH side.

August 2024

Fixed · Global

Due to a maintenance from our infrastructure provider, the MySQL and PostgreSQL DEV clusters of the Montreal (MTL) region will be unavailable on Tuesday, September 3, 2024 starting at 12:00 UTC.

The maintenance is expected to take around 1 hour. During that time, the MTL MySQL and PostgreSQL DEV add-ons will not be available.

This incident will be updated to reflect the maintenance status.

[30/08/2024 15:00 CET] Both cluster are available

MTL: FSBuckets maintenance
Fixed · Global

Due to a hardware maintenance from our provider planned in the next few days, we will need to migrate the FSBucket service of the Montreal (MTL) region on Monday, September 2, 2024 starting at 08:00 UTC.

The maintenance is expected to take less than 1 hour. During that time, the FSBucket service will be read-only. Write operations will be denied. Read operations will continue to work as expected.

All applications linked to an FSBucket add-on on the Montreal region will be redeployed so they can reconnect to the server with read/write rights.

This incident will be updated to reflect the maintenance status.

EDIT 2024-09-02 08:08 UTC: The maintenance is starting. FSBucket are now read-only.

EDIT 2024-09-02 08:24 UTC: Applications are redeployed and should now be able to access their FSBucket.

EDIT 2024-09-02 09:10 UTC: All applications have been redeployed since 08:40 UTC and the maintenance is over. We are still having an issue with the web interface, we are looking into it.

EDIT 2024-09-02 12:16 UTC: The web interface issue has been fixed.

Fixed · Global

Due to a hardware maintenance from our provider planned in the next few days, we will migrate the Git repositories service of the Montreal (MTL) region on Friday, August 30, 2024 starting at 08:00 UTC.

The maintenance is expected to take less than 1 hour. During that time, the Git repositories service will be read-only. Git push operations will be denied. Pull operations will continue to work as expected.

This incident will be updated to reflect the maintenance status.

EDIT 2024-08-30 08:30 UTC: The maintenance is now over. Applications Git deployment URL have changed from push-n1-mtl-clevercloud-customers.services.clever-cloud.com to push-n2-mtl-clevercloud-customers.services.clever-cloud.com. SSH identity should be the same. Using the old domain will keep working for backward compatibility.

Fixed · Infrastructure · Global

A hypervisor is not responding. A VM seems to be stealing all the cpu.

We are force rebooting this hypervisor.

21:09 status: the server refuses to reboot. We asked the OVHCloud support for help.

A technician is having a look at that server. We are waiting for the result of their analysis.

21:33 status: the technician came back to us and signaled a hardware issue. We are waiting for further update and actions.

2024-08-27 07:15 : OVHCloud support finished replacing the motherboard and give us back the server. It fails to reboot outside of rescue. While some are working on getting the kernel to boot, others are moving all the data outside to restore the impacted services for our customers.

09:50 : all services are back up and running for our customers.