Start Your Search Here

Job Search

genesis networks pte ltd

Singapore / Global

AI Infrastructure Engineer

Job Description

1. Compute & ClusterManagement

Architect, configure, and maintain high-density multi-GPU compute clusters (e.g., NVIDIA HGX/DGX architectures).

Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.

Monitor GPU health, telemetry, utilization, and thermals minimize idle compute time and prevent single-node bottlenecks.

2. High-Performance Networking& Storage

Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).

Configure and scale high-throughput parallel file systems and object storage (e.g., Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.

3. Automation &Infrastructure as Code (IaC)

Build and manage automated deployment pipelines using

Terraform, Ansible, Helm, or Pulumi

.

Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.

4. Operations, Observability& Performance

Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).

Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.

Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.

Qualifications &Requirements

Technical Competencies

Operating Systems:

Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).

Accelerated Compute:

Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.

Orchestration & Workload Scheduling:

Hands-on experience with

Kubernetes

(GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).

High-Speed Networking:

Proven experience with RDMA (RoCE v2 / InfiniBand), PFC (Priority Flow Control), and ECN configurations.

Storage Systems:

Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.

Automation:

Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).

Experience & Education

Bachelor's Degree

in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.

3-6+ years

of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.

Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).

Apply Now

Similar Opportunities

View all jobs