Incident History
Full history of incidents.
May 2022
AccessLogs/Metrics are experiencing issues
EDIT 23:41 UTC - Issue has been identified and we are consuming the lag.
EDIT 07:28 UTC - Lag has been consumed .
EDIT 07:30 UTC - Fixed.
Logs are experiencing issues
23:55 UTC - Issues has been identified and we are consuming the lag.
00:19 UTC - lag has been consumed.
00:20 UTC - Fixed.
Metrics/AccessLogs are experiencing issues.
EDIT 09:11 UTC - Metrics/AccessLogs are catching up their lag.
EDIT 16:34 UTC - Fixed.
Logs and drains systems are experiencing issues. We are working on it.
EDIT 09:06 UTC - The logs are catching up.
EDIT 11:15 UTC - Fixed.
Deployment components are experiencing issues to due deployment lag triggered by the Core API issues.
EDIT 08:00 UTC - We have identified ongoing issues.
EDIT 08:02 UTC - New deployments are currently disabled to reduce the impact on our infrastructures. We will reactivate them when the queued ones will be deployed.
EDIT 08:45 UTC - Deployments are still flaky, we are working to resolve the issues.
EDIT 09:08 UTC - Deployments queue is catching up. When it ends, we will redeploy a part of the PAR zone to ensure deployments are monitoring are consistent.
EDIT 09:25 UTC - The mentioned deployments are running.
EDIT 11:16 UTC - We are about at 75% of the deployments completed.
EDIT 12:06 UTC - Finished and fixed.
We are investigating issues with our Core API.
EDIT 06:34 UTC - Our orchestrator is impacted and the deployments are experiencing issues.
EDIT 06:44 UTC - Core API is fixed.
EDIT 08:34 UTC - We are experiencing issues affecting console, cli. We are investigating.
EDIT 08:45 UTC - Core API is fixed.
Clever Cloud API behind the domain name api.clever-cloud.com got some slow-downs
We found out that there is an issue with a shard of our indexes. Some metrics may be unavailable during the reloading period.
Metrics and access logs are currently having ingestion and query issues. We are working on it.
EDIT 12:11 UTC: Metrics and logs are now accessible again. Sorry for the inconvenience.
Metrics and access logs are currently having ingestion and query issues. We are working on it.
EDIT 10:56 UTC: Metrics and logs are now accessible again. Sorry for the inconvenience.
April 2022
Live logs are unavailable EDIT: an internal service was unreachable, Live Logs system is now fully operational
A mongodb node was unreachable. This node is now fixed
Metrics and access logs are currently having ingestion and query issues. We are working on it.
EDIT: Ingestion fixed, query almost restored
Logs pipeline components lost their connection EDIT: Connection issue fixed, we are consuming logs queue lag. we lost almost30 min of logs. EDIT: Lag consumed.
We are facing an issue with the indexes which result on some metrics and access logs unavailability at egress-level. EDIT: all indexes has been rebooted
Metrics and access logs are currently having ingestion issue. We are working on it.
EDIT 13:20 UTC: The issue has been fixed. Some metrics data points have been lost. Access logs are being queued for ingestion again.
Some of API calls might return a 504 error. The source cause has been found and we are working to restore the service.
EDIT 17:20 UTC: The service has been fully restored. Sorry for the inconvenience.
We have identified issues affecting logs and drains. We are working on it.
EDIT 06:45 UTC: fixed.
EDIT 07:22 UTC: we have identified another issue.
EDIT 09:45 UTC: fixed.
We are currently experiencing instabilities with our main API. We are looking into it.
EDIT 15:12 UTC: This seems to be back to normal. We did not find the root cause but we keep looking. Some actions may have failed like deployments, git push or accessing the dashboard / using the CLI in general
EDIT 17:34 UTC: We still see some instabilities, resulting in various longer queries or even errors from some services that fail to contact our API. We are still working on identifying the root cause.
EDIT 20:34 UTC: We didn't see any more instabilities since the latest status update. We'll continue to monitor the activity in the next couples of days.
Access logs currently have a few hours of ingestion delay. It is currently being resolved and the delay should be back to normal in a few hours. This impacts the retrieval of access logs using the CLI or the API. Also, the various console dashboards (status codes, requests per hours, ...) are impacted and might display out of sync data. Sorry for the inconvenience.
EDIT 20:43 UTC: The delay has now resolved, you should now be able to query the access logs using the CLI or API.