dadaconsultants pte. ltd.
Alexandra, Singapore Country / Global
Alexandra, Singapore Country / Global
About the Role
Our client is a world-leading technology company operating at the intersection of high-performance computing, Bitcoin mining, and AI cloud services, with a globally distributed infrastructure footprint spanning multiple continents and a multi-gigawatt energy portfolio. Headquartered in Singapore, they are scaling rapidly and investing heavily in next-generation datacenter and cloud capabilities. This role sits at the deepest technical tier of the operations centre, owning the most complex customer escalations and platform incidents as the critical bridge between front-line support and Engineering/SRE teams. It is an exceptional opportunity for a seasoned infrastructure engineer to drive permanent, systemic improvements rather than just resolving individual tickets. You will directly shape the reliability and quality of a cloud platform serving high-demand AI and HPC workloads.
Key Responsibilities
Own complex, escalated customer issues and platform incidents end-to-end, driving each case through to full resolution
Conduct deep-dive troubleshooting across GPU compute , networking/SDN, storage, drivers, control plane, and billing systems
Serve as the primary liaison to SRE, Compute, and R&D teams - leading root-cause analysis and ensuring permanent fixes are implemented
Lead and support incident response and post-incident reviews, translating findings into improved runbooks and monitoring enhancements
Identify recurring escalation patterns and convert them into lasting platform or process improvements
Mentor L1 and L2 support engineers, raising escalation quality and expanding the team knowledge base
Participate in an on-call escalation rotation, providing incident leadership during critical platform events
Collaborate cross-functionally with engineering and product teams to feed operational learnings back into the platform roadmap
Requirements
Minimum 5 years of hands-on experience in cloud infrastructure technical support, escalation engineering, or SRE-adjacent roles
Strong practical proficiency in Linux, networking, and cloud infrastructure experience with GPU/CUDA/HPC environments (is a bonus)
Demonstrated track record of resolving complex production incidents and conducting thorough root-cause analysis
Excellent written English communication skills, with the ability to work effectively across technical and non-technical stakeholders
Comfortable with on-call responsibilities and experienced in taking ownership during high-pressure incident scenarios
Familiarity with SDN, distributed storage systems, or control plane architecture (is a bonus)
If you are passionate about technology and meet the above requirements, please don't hesitate to apply. Please note that only shortlisted candidates will be contacted. Appreciate your understanding. Data provided is for recruitment purposes only.
Dada Consultants Pte Ltd
Website:
EA License No.: 18S9037
Business Registration Number: 201735941W
Ubi, Singapore Country / Global
Orchard, Singapore Country / Global
Orchard Road, Singapore / Global
Cross Street, Singapore / Global
Cecil, Singapore Country / Global
Cecil Street, Singapore / Global