Platform Engineer (Ceph Storage Specialist) - #1678
Must-have:PythonNode.jsKubernetesDevOps
Platform Engineer (Ceph Storage Specialist)
[What you will be working on]
We are looking for a hands-on storage engineer to operate and maintain the Ceph-based storage services supporting our on-premises Kubernetes and OpenShift platform.
The role focuses on operating and improving production storage services, including monitoring, maintenance, upgrades, capacity management, hardware replacement, performance troubleshooting, and failure recovery. We are not developing a new distributed-storage solution. Candidates with experience developing or operating distributed-storage platforms are encouraged to apply.
You will work closely with the Kubernetes platform, network, infrastructure, and application teams to keep storage services reliable and supportable.
What to Expect
- Operate and maintain Ceph and OpenShift Data Foundation environments.
- Monitor cluster health, capacity, latency, throughput, placement groups, and device status.
- Manage OSD and node replacement, recovery, rebalancing, backfill, scrubbing, and routine maintenance.
- Diagnose degraded placement groups, slow operations, quorum issues, device failures, and network-related storage problems.
- Plan and perform Ceph/ODF upgrades, expansion, patching, and configuration changes.
- Support block and shared-file storage through ODF, Ceph CSI, RBD, and CephFS.
- Troubleshoot persistent-volume provisioning, attachment, mounting, expansion, and performance issues.
- Maintain dashboards, alerts, runbooks, capacity plans, and recovery procedures.
- Test recovery from disk, node, service, and network failures.
- Coordinate server, disk, firmware, and network maintenance with infrastructure teams.
- Participate in production incident response and root-cause analysis.
[What we are looking for]
- Hands-on experience with Ceph
- Experience with storage monitoring, capacity planning, upgrades, expansion, and component replacement.
- Experience benchmarking storage workloads and analysing performance bottlenecks.
- Ability to diagnose storage issues across software, Linux, networks, and physical hardware.
- Knowledge of HDD, SSD, NVMe, HBA, firmware, and storage-network considerations.
- Automation experience using Ansible, Python, shell, or similar tools.
- Backup, snapshots, replication, and disaster-recovery operations.