[PAR] Ceph instabillity during maintainance
Resolved:
A maintenance operation is currently in progress on the Ceph infrastructure in par9, including the addition of new nodes to the cluster.
During this operation, part of the Ceph stack became unstable, which may have caused intermittent errors on Cellar / object storage access, including 503 responses observed by monitoring.
Active changes have been paused while the team investigates why some components are going down and coming back online. The cluster remains under close monitoring until it returns to a nominal state.
Updates
The maintainance incident was linked to a network misconfiguration on nodes added to the cluster which lead to cycle of up and down of the osds on the new servers and disruption of the ceph state.
We have finished the maintainance operation and everything is working properly