Incident History
Full history of incidents.
June 2024
Some logs drains are not correctly sent to their target. The issue seems to have started since 2024-06-21 10:30 UTC+2. We identified the issue and are working towards a fix.
EDIT 16:53 UTC+2: Logs drains should now be fully functional since 14:13 UTC+2 and are stable since then. If you still have missing logs from your drains, please open a ticket and we will investigate it further.
The monitoring has detected a cut in network traffic for the Paris datacenters, we are investigating the issue.
EDIT 06:05 UTC : The network traffic is come back as its casual rate. We have seen a cut outside Clever Cloud network, we are investigating why.
EDIT 06:15 UTC : We have seen a second cut, we have identified that a network provider is doing maintenance operation which seems to be the cause.
EDIT 07:21 UTC : We have seen a third cut.
EDIT 07:40 UTC : We have contacted our network provider and confirmed that cuts are coming from the maintenance. For now, we are aware of 5 cuts due to the maintenance.
EDIT 08:10 UTC : We are not expecting more network cuts as the maintenance window is over, but we are watching.
At 20:30 UTC, our monitoring registered a wave of network reconnections and downtime of IPSec tunnels for a few minutes.
We checked all the tunnels and restarted the ones that did not restart automatically. We checked the load balancers and did not see anything strange except the spike on reconnections.
After investigating, our probes revealed a very low rate of packets from the internet for 5 minutes.
We are encountering delivery issues with the Logs Drains platform. We are currently investigating the issue. Logs drains delivery may be delayed until this issue is resolved.
EDIT 09:30 UTC: We may have found the origin of the issue and implement a fix. We are monitoring the fix. Currently, logs drain are delivered without delay.
EDIT 14:09 UTC: The situation is now stable. Incident is closed.
Our main API was unavailable for a few minutes between 13:20 UTC until 13:23 UTC. We are looking into it. Deployments started during that period may be impacted.
EDIT 2024-06-04 13:32 UTC: The root cause has been found and fixed. Deployments that were started during that period may have failed. You should be able to retry them. Please contact our support team if you still face any issues.
We have lag on the ingestion pipeline, we are investigating the issue.
EDIT 20:00 UTC : We are still investigating the issue.
EDIT 22:00 UTC : We are still seeing lag on the ingestion pipeline, we have found that there is a bottleneck on the offload process of the tiered storage of pulsar that we are fixed, but we now need to wait pulsar to finish its offload process.
EDIT 2024-06-07 15:00 UTC : The pulsar cluster have finished the offload tasks and we are recovering the lag starting yesterday around 16h utc.
EDIT 2024-06-10 08:00 UTC : We have finsihed to recovers the access logs lag since saturday 14:00 utc.
May 2024
We are experiencing some network issues due to our infrastructure provider network misconfiguration. We are currently observing high latency and packet loss when trying to reach machines in Singapore. A ticket is being created
EDIT 10:30AM UTC : A ticket has been opened with our infrastructure provider
EDIT 12:45PM UTC : Network seems more stable since 12:20PM UTC
EDIT 17:00PM UTC : The ticket with our infrastructure provider has been closed
Maintenance Window: 2024-05-27T09:00:00Z - 2024-05-29T20:00:00Z (UTC)
Scope:
- We will roll out software updates on PAR region and dedicated load balancers
Expected Impact:
- Brief disconnections or connection drops during the upgrade process.
- Potential minor performance fluctuations.
Additional Information:
- Please report any issues with a method for reproducing the problem (e.g., curl command for application load balancer issues).
EDIT 2024-05-28 13:15 UTC : We are beginning the maintenance. We are starting with cleverapps.io.
EDIT 14:10 UTC : We have updated cleverapps.io load balancers, we are updating the par region one.
EDIT 15:30 UTC : We have updated three of nine load balancers of par region, the updates are still running
EDIT 16:30 UTC : We have update six of nine load balancers of par region, the are still running for the last ones.
EDIT 17:00 UTC : We have finished to update the par load balancer, we will do the dedicated one starting tomorrow
EDIT 2024-05-28 08:00 UTC : We have seen an increase of tls error on par region, we have rollback 8 of 9 instances of this load balancer on the previous version which is not affected. We keep one instance with the issue the time to dig and found the the root cause.
EDIT 12:30 UTC : We have found the issue and written a patch, we are releasing it and then we will deploy the new version. The issue was limited to services under *.services.clever-cloud.com certificate only.
EDIT 13:30 UTC : We have deployed the new release on par region, we will start very soon with others regions and cleverapps.io
EDIT 14:50 UTC: We have deployed the new release on every region including dedicated ones and cleverapps.io. We will begin the dedicated load balancers very soon.
EDIT 17:20 UTC : We are finishing the last dedicated load balancers for today and we will terminate the others tomorrow.
We had an issue with the compute pipeline of telemetry from access logs which did not perform the computation since the end of the last week. We have fixed the issue, but we could not recover missing computations.
We identified an increase of errors on read queries, writing isn't impacted.
Maintenance Window: 2024-05-22T09:00:00Z - 2024-05-24T20:00:00Z (UTC)
Scope:
- We will roll out software updates on every region of Clever Cloud for all application load balancers
Expected Impact:
- Brief disconnections or connection drops during the upgrade process.
- Potential minor performance fluctuations.
Additional Information:
- Please report any issues with a method for reproducing the problem (e.g., curl command for application load balancer issues).
EDIT 13:06 UTC : We are beginning the rolling of the WSW region.
EDIT 15:00 UTC : We have updated the WSW region. We are beginning the rolling of SGP, MTL and SYD regions.
EDIT 16:25 UTC : We have updated the MTL region.
EDIT 17:00 UTC : We have updated the SYD and SGP region. Next region tomorrow
EDIT D+1 08:10 UTC : We will begin the update of MEA and GRA-HDS.
EDIT D+1 09:20 UTC : We have finished the update of GRA-HDS, we are beginning RBX and RBX-HDS. The rolling update of MEA is still running.
EDIT D+1 11:15 UTC : We have finished the update of MEA, RBX and RBX-HDS.
EDIT D+1 15:00 UTC : We are beginning the update of SCW and dedicated regions.
EDIT D+2 13:00 UTC :We have finished the update of SCW and dedicated regions. We will perform the update of the PAR regions and dedicated application load balancers on a new status here : https://www.clevercloudstatus.com/incident/855
An add-on reverse proxy was unreachable from 14:23 UTC+2 to 14:27 UTC+2. During that time, connections to add-on services might have timed out or failed with various errors.
The issue has been resolved.
We identified an increase of errors on read queries, writing isn't impacted.
UTC 08:53 : Read queries have been disabled in order to solve the issue UTC 09:23 : Read queries are now back to normal.
The following domains seem to have been reported as unsafe to Microsoft Defender SmartScreen:
- cellar-c2.services.clever-cloud.com
- cellar-fr-north-hds-c1.services.clever-cloud.com
- cellar-fr-north-c1.services.clever-cloud.com
This means that Microsoft Edge users might have issues downloading files stored on our Cellar services.
We are currently in discussion with them to get the domains unblocked. In the meantime, do not hesitate to click the "This domain is safe" link in the Defender screen.
EDIT 2024-05-07
We have setup a workaround. To avoid this workaround to be used by malicious users, we won't disclose it here.
Please come to us if you need to use it.
EDIT 2024-06-28
The domains were un-flagged a few days ago. The incident is now over.
(Times are in UTC)
At 2024-05-04 23:31 a database load balancer lost its network routes. The alert about that was set as low priority and did not wake up the on-call agent. At 01:23, another service failed because of that load-balancer issue. This time, the failure triggered a high priority alert.
The on-call agent investigated the issue and saw that the load-balancer was responsible for the other service's failure. They fixed the network issue. Every impacted service got back online around 01:45.
Clever Cloud's PAR region has 8 of those load balancers. Only the services that were trying to connect to this one got downtime. Some customers applications redeployed themselves and connected to another one, quickly fixing the issue.
On 2024-05-06, we made the first alert a high priority one. It should already have been high priority. We also made sure that every other "load balancer is unreachable" alerts were high priority ones.
We are observing timeouts and errors on api.clever-cloud.com. We are investigating the issue.
EDIT 16:00 UTC : We have found an issue, we are patching it and redeploying the api.
EDIT 16:10 UTC : We have deployed a new version of the api
EDIT 16:20 UTC : The issue seems to be solved, we are keeping a eye on it
EDIT 16:30 UTC : The issue is solved we did not observe errors and timeouts
April 2024
Some emails issued by the heptapod service weren't correctly delivered to their recipients the last few days. The underlying issue has been fixed and the mail backlog is currently being processed. Additional monitoring will be put in place to monitor the email queue.
We will update this incident once the backlog is fully processed.
EDIT 2024-04-25 16:00 UTC: The backlog has been fully ingested. The incident is now over.
An operation on the metric cluster is pending which will make it more resilient to spikes and load. It shouldn't impact read queries of metrics, it can generate lag in the writing path.
EDIT UTC 18:29 : Operation is done, services weren't disturbed.
Beginning at 5h00 UTC, we seen a drop in the rate of access logs consumption which seems to be caused to difficulty to produce them. We are investigating the issue. You may see delays to retrieve your access logs.
EDIT 10:30 UTC : We are performing a rolling restart of the underlying pulsar brokers, you may seen disconnection.
EDIT 16:00 UTC : The rolling restart is performed. We still have ingestion issues we will keep investigating
EDIT D+1 08:50 UTC : We have still ingestion issues on few partitions which may be related to an underlying trouble, we are digging into it.
EDIT D+2 14:00 UTC : We have found the underlying issue and solve it, we are consuming the remaining lags.
EDIT D+3 13:00 UTC : We are still consuming the remaining lags, the current eta of full recovery is targeting tomorrow during the night
EDIT D+4 06:00 UTC : We have done consuming the remaining lag.
We are currently experiencing a disruption in our email services due to an unforeseen issue, emails will be delayed until this issue is resolved. Our team is actively working to restore access as quickly as possible. We will keep you updated on our progress and notify you as soon as services are fully operational again.
EDIT 20:04 UTC+2: We are still working on the issue.
EDIT 2024-04-19 12:17 UTC+2: The issue has been fixed, we continue to monitor the situation.