Platform Engineer (Ceph Storage Specialist)
Must-have:PythonNode.jsKubernetesDevOps
Job Description & Requirements
We are seeking a hands-on Platform Engineer specializing in Ceph storage to operate and maintain production storage services supporting on-premises Kubernetes and OpenShift environments.
Key Responsibilities
- Operate and maintain Ceph and OpenShift Data Foundation (ODF) storage environments.
- Monitor cluster health, capacity, latency, throughput, placement groups and device status.
- Manage OSD/node replacement, recovery, rebalancing, backfill and scrubbing activities.
- Troubleshoot degraded PGs, slow operations, quorum issues, device failures and storage-network problems.
- Perform Ceph/ODF upgrades, expansion, patching and configuration changes.
- Support RBD, CephFS, Ceph CSI and persistent volumes in Kubernetes/OpenShift.
- Troubleshoot PV provisioning, attachment, mounting, expansion and performance issues.
- Maintain monitoring dashboards, alerts, runbooks, capacity plans and recovery procedures.
- Perform recovery testing for disk, node, service and network failures.
- Coordinate server, disk, firmware and network maintenance with infrastructure teams.
- Participate in production incidents, troubleshooting and root-cause analysis.
Requirements
- Hands-on experience operating Ceph in production .
- Experience with storage monitoring, capacity planning, upgrades, expansion and hardware/component replacement.
- Strong troubleshooting skills across Ceph, Linux, Kubernetes/OpenShift, networking and physical infrastructure .
- Experience analysing storage performance and benchmarking workloads.
- Knowledge of HDD, SSD, NVMe, HBA, firmware and storage networking .
- Experience with Ansible, Python, Shell scripting or similar automation tools .
- Knowledge of backup, snapshots, replication and disaster-recovery operations.
- Experience with Kubernetes/OpenShift storage, CSI, RBD or CephFS is highly preferred.