digital biz solutions pte. ltd.
Singapore / Global
Singapore / Global
About the Role
We are looking for a hands-on Platform Operations Engineer (Ceph Storage Specialist) to operate and continuously improve Ceph-based storage services that support on-premises Kubernetes and OpenShift platforms for one of our clientele project that is going live soon. This role is operations-focused (production reliability and supportability), covering monitoring, maintenance, upgrades, capacity management, hardware lifecycle, performance troubleshooting, and failure recovery. You will work closely with platform, network, infrastructure, and application teams to keep storage services reliable, scalable, and well-governed.
Key Responsibilities
Operate and maintain production
Ceph
and
OpenShift Data Foundation (ODF)
environments supporting Kubernetes/OpenShift platforms.
Monitor and manage storage health and performance across
capacity, latency, throughput, placement groups (PGs), and device status .
Perform routine operational tasks including
rebalancing, backfill, scrubbing, recovery actions , and general cluster maintenance.
Plan and execute
Ceph/ODF upgrades, patching, expansions, and configuration changes
with minimal service disruption.
Manage
OSD and node replacement , including coordinating server/disk/firmware maintenance with infrastructure teams.
Diagnose and resolve production incidents involving
degraded PGs, slow ops, quorum issues, device failures, and network-related storage problems .
Support Kubernetes/OpenShift storage consumption patterns, including: (
Ceph CSI, RBD (block storage) & CephFS (shared file storage))
Persistent Volume lifecycle troubleshooting (provisioning, attachment, mounting, expansion, and performance)
Maintain and improve operational readiness through
dashboards, alerts, runbooks, capacity plans, and recovery procedures .
Test and validate recovery from
disk, node, service, and network failures , and continuously improve resilience.
Participate in production incident response, post-incident review, and
root-cause analysis (RCA)
drive corrective and preventive actions.
Collaborate with cross-functional stakeholders (platform, network, infrastructure, application teams) to ensure storage services remain reliable and supportable for clientele platforms.
Required Skills and Experience
Hands-on experience operating
Ceph
in production environments (day-2 operations, troubleshooting, maintenance).
Experience operating
OpenShift Data Foundation (ODF)
and supporting storage services for Kubernetes/OpenShift platforms.
Strong experience with storage operations, including
monitoring, capacity planning, upgrades, expansion, and component replacement .
Proven ability to troubleshoot complex storage issues across
software, Linux, networks, and physical hardware .
Working knowledge of storage hardware and performance considerations, including
HDD, SSD, NVMe, HBA, firmware , and storage-network behaviour.
Experience supporting block and shared-file storage using
Ceph CSI, RBD, and CephFS , including PV-related troubleshooting.
Experience benchmarking storage workloads and analysing
performance bottlenecks
(latency/throughput, saturation points, noisy neighbour symptoms).
Experience in incident handling and operational discipline, including
on-call/incident response , triage, and RCA practices.
Automation/scripting capability using
Ansible, Python, shell , or equivalent tools to reduce manual operations and improve consistency.
Strong documentation skills to maintain
runbooks, recovery procedures, and operational standards .
Preferred Skills
Experience with
backup, snapshots, replication, and disaster recovery (DR)
operations for Ceph/ODF-backed platforms.
Experience building or improving observability for storage platforms (dashboards, alert tuning, actionable SLO/SLA signals).
Familiarity with platform/SRE practices such as reliability engineering, change management, and continuous improvement in production environments.
Exposure to Kubernetes/OpenShift platform operations beyond storage (helpful for cross-team troubleshooting and incident coordination).
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global