Incident History
Full history of incidents.
June 2025
Some deployments are currently blocked, we are investigating the issue
We are investigating various issues on the Paris region.
Metrics service is currently under heavy load and has issues to hold read/write operations. We are currently investigating and fixing the issue. Lag on metrics can be seen (few minutes at most), no data is lost.
15:14 UTC: Situation is back to normal
Following friday’s incident, FS bucket creation in several regions were prevented by a sporadic network issue.
This also prevented deployments for new PHP applications that did not disable the automatic FS bucket.
This is now resolved.
Between 12:00 UTC and 13:40 UTC all the load balancers stopped consuming orders. The monitoring alerted in a weird way, which led to a slower on-call response than expected.
The situation has been resolved
We identified availability issues on newly created add-on on the gra-hds zone. We are investigating
A maintenance led to a disruption which can prevent to create addons and applications. Panels in the console are impacted as well.
Applications and addons are still running, only the console is impacted.
12:22 UTC: Problem is identified, our on-call team is investigating and deploy fixes
13:29 UTC: Problem is fixed
A hypervisor on RBX is not responding.
All applications on it have been redeployed. The databases services running on it are unreachable.
We are rebooting the machine and investigating
Our Paris region had availability issues from certain networks. The issue was coming from one of our network transit provider and started around 22:40 UTC+2. We stopped the peering session with our transit provider at 23:50 UTC+2 and the situation is now back to normal.
Traffic incoming or outgoing through this transit provider might have been lost or had increased latencies during this incident.
We continue to monitor the situation.
08:19 UTC: Our ingresses & egresses received no more requests from customers
08:25 UTC: We re-established the connectivity with sampling on data. 08:35 UTC: Incident solved
We are investigating an issue which prevents from creating new accounts or organisations.
May 2025
A hypervisor crashed in our PAR region. It rebooted itself 3 minutes later.
We are currently checking and restarting all the services it holds.
We are currently experiencing an unexpected reboot of a hypervisor in our Paris region data center. This incident has led to temporary disruptions in service for some of our hosted services on that hypervisor. We are investigating the cause of the reboot and working to restore normal operations as quickly as possible. We are prioritizing the recovery of critical services and applications.
At 15:15 UTC a database upgrade made a bug in the API visible. Applications could not be created when following the console’s application creation screen.
The bug was found and fixed.
We are investigating an outage impacting service availability for heptapod.host.
The deployment API is currently unavailable, leading to crashed applications to remain inaccessible until the issue is resolved. However, the Clever Cloud API is functioning properly, while the web console is currently down. On-call teams are diligently working on resolving the issue.
EDIT 16:12 UTC: We are restarting the deployment stack.
EDIT 16:48 UTC: The deployments have resumed, and the team is closely monitoring this system in particular.
EDIT 16:59 UTC: All systems appear to be functioning properly. The team will continue monitoring the situation.
EDIT 17:48 UTC: All systems continue to function properly. Only some instances initialization on one of our PAR AZ is still triggering occasional errors. We continue to investigate it.
April 2025
A hypervisor has crashed and is currently rebooting. It has an impact on deployments in par6 availability zone. All the databases located on this machine are currently down.
The SSH gateway fails to setup the temporary keys on the VMs. This comes from a certificate issue with the underlying AMQP cluster.
We are investigaging the source cause.
We are monitoring our mailing system for our status page. Some emails seem not to be properly delivered
The GRAHDS region is currently unreachable. This seems to be a problem in OVHcloud network as we can't reach it from the public internet or from other OVHcloud regions.