Incident History
Full history of incidents.
May 2018
We have started a maintenance operation on a component of the Metrics cluster. This operation takes more time than expected.
Until it's over, Metrics are not available. Metrics agents on scalers should push the data when the service is back.
EDIT 15:14 UTC: Metrics are back since 15:12 UTC
April 2018
A dedicated add-ons reverse proxy stopped accepting new connections at 15:28 UTC and was restarted at 15:31:30 UTC.
Traffic was back to normal at 15:32:00 UTC.
Logs drains are currently stopped, we are working on fixing this issue.
Due to a network issue, deployments are not working properly. Also, the state of the applications might be displayed wrong (grey disc instead of green one) in the console.
March 2018
Cellar is having network issues on some node. Some requests are failing, both requests to GET resources as requests to send resources.
We are investigating the problem
EDIT 19:35 UTC: The problem seems to be gone.. It may be due to a maintenance operation made on the Cellar cluster which shouldn't have caused this. This maintenance has been done multiples times without problems. We will keep an eye on the cluster when this maintenance starts again, probably tomorrow.
Multiple reports are indicating there is a network slow down for some clients. We are investigating the issue. Applications may take higher time than usual to respond
EDIT 15:10 UTC: The source of the problem is one of our customers receiving a DDoS on its application. While the infrastructure can handle such load, we detected a problem with the configuration of our reverse proxies which doesn't allow us to correctly handle the load of this DDoS. We are looking at how we can improve that. In the meantime, traffic targetting that customer's application has been blocked.
EDIT 16:45 UTC: Most of the traffic is filtered. We will continue watch the issue in the following hours
A dedicated addons reverse proxy is refusing new connections. It is being restarted.
EDIT 11:49 UTC: Incident over since 11:45 UTC
Real-time log delivery is affected by an outage on our message broker. Log drains are affected as well. Logs are still archived.
EDIT 17:03 UTC: Real-time delivery is back since 16:50 UTC
February 2018
Our monitoring system has detected network connectivity issues. Issues were caused by a network configuration inconsistency, they are solved.
NodeJS applications are failing to deploy because of the missing nomnom module. We are investigating the issue.
EDIT 10:53 UTC: You can create the following environment variable for a temporary workaround: CC_PRE_RUN_HOOK=npm install nomnom@1.8.1 -g
EDIT 11:33 UTC: A fix has been made and the new image version is now deploying on our servers.
EDIT 12:33 UTC: The new image is now live. All NodeJS applications will be redeployed to avoid using a now broken image.
The metrics data cluster is under unusual load. Metrics display is currently unavailable, but metrics are still collected.
EDIT 17:35 UTC: Service is back to normal and collected metrics have all been correctly persisted.
The proxy is being restarted. Some add-ons may be unreachable until it's done.
EDIT 15:42 UTC: Incident over since 15:40 UTC.
The log storage cluster is experiencing network issues. We are working on it. In the meantime, only realtime logs are available.
January 2018
The proxy is being restarted. Some add-ons may be unreachable until it's done
EDIT 16:41 UTC: the proxy has been successfully restarted. Add-ons should be reachable again. Applications not supporting the loss of an established connection will be redeployed. We continue to monitor the proxy.
EDIT 17:30 UTC: the incident is now over
A redis cluster was down and is restarting
EDIT 20:17:00 UTC: The cluster has been restarted, impacted applications have been redeployed. The incident is over
PostgreSQL addon dashboards will be unavailable for about 15 minutes starting on 2018-01-25 at 12:30 UTC
EDIT: Delayed to 12:50 UTC
EDIT 12:50 UTC: Will start in a few seconds
EDIT 13:07 UTC: Maintenance over. If you encounter an issue, please tell us.
Logs are currently unavailable. We are working on restoring them. All logs sent in the last 30 minutes won't be stored.
EDIT 03:15 UTC: Logs are back again
The MongoDB shared cluster needs to be upgraded to have more resources.
Performance issues and or partial outage are to be expected. We will try to keep them as low as possible.
The maintenance starts at 22:00 UTC
EDIT 02:00 UTC: the maintenance is now over
An addon reverse proxy is restarting, connections are dropped and impacted applications will be redeployed
EDIT 20:45:00 UTC: The reverse proxy took ~1 minute to restart. It is now restarted
EDIT 20:48:00 UTC: Impacted applications were redeployed as expected. The incident is now over and all add-ons are now reachable again
All deployments from around 15:40 UTC might be shown in a FAILED state, even though they were successful. It's just a matter of display and the instances, if correctly deployed, are put into production.
The Activity pane (Console), clever status (cli) and the API endpoint /applications/<app>/deployments incorrectly report the deployment status.
Notifications (slack webhooks, mails) correctly report the deployment status (failed or successful) and can be trusted.
EDIT 21:48 UTC: It should now be fixed. Deployments with the "FAILED" state will keep their broken state.