Start Your Search Here

Job Search

runsun service pte. ltd.

Singapore / Global

Hardware Engineer

Job Description

We are seeking an experienced AI Hardware Engineer to support the design, deployment, validation, and troubleshooting of AI training clusters, GPU servers, networking, and storage infrastructure. The ideal candidate should have strong expertise in server hardware, GPU platforms, high-speed networking, and data center infrastructure to support large-scale AI/HPC environments.

Key Responsibilities AI Server Hardware Management

Deploy, validate, and maintain AI GPU servers

Perform hardware diagnostics and component replacement

Analyze system logs, BMC logs, and hardware alerts

Manage server hardware lifecycle.

GPU Platform Support

Deploy and validate NVIDIA GPU platforms

Troubleshoot GPU-related

Perform GPU benchmarking and stress testing

Support CUDA, NCCL, and GPU fabric troubleshooting.

AI Cluster Deployment & Validation

Participate in AI/HPC cluster deployment

Execute cluster hardware qualification testing

Produce validation reports and documentation.

Network & Storage Support

Configure and maintain high-speed networking:

Support distributed storage systems:

Assist with performance analysis and troubleshooting.

Automation & Tool Development

Develop automation scripts for:

Hardware health checks

Cluster validation

Deployment automation

Log collection

Build tools for testing and operations.

Good communication, teamwork, and ownership mindset.

Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.

Required Qualifications

Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.

Hardware

Strong knowledge of x86 server architecture

Familiar with Intel, AMD, and NVIDIA Grace CPU platforms

Experience with:

HGX

DGX

GB200 NVL72

GB300 NVL72

Knowledge of BMC/IPMI management.

GPU & AI Platform

Experience with NVIDIA GPU products H100, H200, B200, B300

Familiar with: CUDA ,NCCL ,NV Link ,NV Switch and GPU Direct RDMA

Linux

Strong Linux administration skills (Ubuntu, Rocky Linux)

Proficient in: Shell ,Python , Bash

Capable of independent troubleshooting.

Networking: Strong understanding of: TCP/IP , VLAN , BGP ,OSPF ,RDMA , InfiniBand and RoCE

Preferred Qualities

Experience operating AI training clusters Kubernetes experience Slurm administration

PXE deployment experience GPU Fabric Manager expertise

Experience with hyperscale AI datacenter deployments.

Apply Now

Similar Opportunities

View all jobs