Incident History
Full history of incidents.
August 2023
Between 13:51 UTC and 14:12 UTC, some requests may have failed to establish a TLS connection to the Cellar service.
The issue has been identified and has been fixed.
We detect that the ip 212.129.27.183 is unreachable, we have identified the root cause and we are waiting for the feedback of scaleway cloud provider.
EDIT 12:39 UTC : The ip address is reachable
The storage layer of metrics and access logs has lost some data nodes. We are fixing the issue
EDIT 09:18 UTC : We are recovering from the events and consuming the lags. The storage layer is now operational
July 2023
We are detecting some errors on our reverse proxies, your apps may not be reachable. We are working on it.
EDIT 22:48 PM UTC: all reverse proxies are now working properly
An issue with the control plane triggered some issues when ordering or migrating MySQL add-ons.
Edit 12:50 UTC: Control plane has recovered, everything is now OK
We are currently looking into an issue regarding applications deployments. They may be able to start but may never complete.
EDIT 15:05 UTC: The issue appears to be limited to the Paris zone
EDIT 15:20 UTC: A counter measure has been deployed to mitigate issues. Deployments are now scheduled as expected. Some errors may still appear in your Logs. We're processing stuck deployments, but you may cancel or start a new one if you want to prioritize your deployment.
We have asked our provider to transfer the domain name cleverapps.io. The transfer ends at 12:30 UTC and we saw that records are missing or have not the right value.
EDIT 15:00 UTC : we have found that NS records and SOA records was not good, we have updated it. EDIT 16:00 UTC: everything is back to normal.
Following the yesterday deployment, we had issues with http and tcp redirections which cause infinite loop and timeouts. We are investigating the issue.
EDIT 09:00 UTC The issue was found and fixed
Some apps are not availables
21h11: only apps with redirect_https enabled are impacted
21h56: we rollback to the old cleverapps loadbalancers
20h15: We have lost our hypervisors on SYD region 20h30: Our infrastructure provider on SYD lost its connectivity 21h00: hypervisors are back online
We are currently experiencing issues on reverse proxies of the JED region. We are investigating them.
EDIT 16:44 UTC: The root cause has been identified and a fix has been applied. We are monitoring the results.
EDIT 16:50 UTC: The service is now operational.
An FSBucket server is currently being investigated for connection timeouts when mounting buckets. The problem has been partially identified and a first fix has been applied. Additional steps will be taken shortly to make sure everything is working as intended.
EDIT 10:49 UTC: The underlying issue has been fixed. Some applications may have had troubles mounting FSBuckets, writing or reading files stored on that server between 08:50 UTC and 10:25 UTC. Impacted applications are currently being redeployed out of caution (most of them successfully reconnected to the server after the fix has been issued).
June 2023
We are detecting some errors on our storage layer responsible for storing metrics and access logs data. Queries were unavailable.
Edit 14:10 UTC: query is re-open
We continue to investigate.
The monitoring system has detected that an hypervisor is unreachable. We are investigating.
EDIT 16:27 UTC: The hypervisor took some time to reboot but it is now up and running. We are making sure services are working fine following this incident.
EDIT 17:10 UTC: The incident is now over. The underlying problem has been identified but the hypervisor is currently in the upgrade queue.
An hypervisor on the Paris region needs to be rebooted due to a kernel issue. The reboot will take place tonight (June 21, 2023) at 18:00 UTC. Services on that hypervisor are already migrated apart for a few of them. Impacted customers will shortly receive an email with more details.
EDIT 18:14 UTC: The maintenance is starting
EDIT 22:00 UTC: The maintenance is now over
An hypervisor on the Paris region needs to be rebooted due to a kernel issue. The reboot will take place tonight (June 21, 2023) at 20:00 UTC. Services on that hypervisor will be migrated starting at 18:00 UTC. Impacted users will shortly receive an email with more details.
EDIT 18:13 UTC: The maintenance is starting
EDIT 23:11 UTC: The maintenance is now over
A deployment issue has been identified, we are working on a fix.
EDIT 20:43 UTC - fixed.
We are detecting some errors on our storage layer responsible for storing metrics and access logs data. We are investigating.
Edit 04:58 PM UTC: A storage node had a hardware issue, it has been rebooted.
We will start a maintenance this Tuesday designed to improve performance on our storage layer for metrics and access-logs. During the maintenance, you may not see latest datapoints and access-logs.
Maintenance will start 20 of June, at 03:30 PM UTC.
Edit 03:45 PM UTC: maintenance is starting.
Edit 04:58 PM UTC: maintenance is over.
MySQL add-on API started to timeout while trying to create add-ons. Currently created add-ons still work, though.
We are investigating the issue.
EDIT 09:00 PM UTC: the root cause has been corrected.