Start Your Search Here

Job Search

lightwheel

Singapore / Global

Cloud Infrastructure Engineer

Job Description

About the Role

The Cloud Infrastructure Engineer (Distributed Task Scheduling) will be responsible for building and optimizing the distributed computing platform that powers human data and simulation data production.

Responsibilities

Build the distributed computing platform that powers human data and simulation data production.

Develop cross-region compute resource scheduling, task queues, quotas, priorities, and elastic scaling capabilities.

Break down sequential tasks into parallelizable, isolated, retryable, and observable compute units.

Support batch, micro-batch, and streaming execution modes; strengthen fault recovery and capacity governance.

Continuously optimize task throughput, queueing latency, failure rate, GPU utilization, and unit compute cost.

Provide clear, observable task interfaces for both systems and agents.

Qualifications

3+ years of experience in distributed computing, task scheduling, data platforms, or infrastructure.

Bachelor's degree or above preferred; outstanding candidates may be considered regardless of degree.

Solid understanding of task scheduling, queues, concurrency, resource pools, isolation, and fault recovery, with hands-on production experience.

Experience with cloud computing, GPU/compute accelerators, or large-scale asynchronous task platforms.

Ability to identify parallelization boundaries, data dependencies, and resource isolation boundaries in complex processing pipelines.

Hands-on experience designing, implementing, debugging, and operating production systems.

Proficient with AI-assisted development tools such as Cursor, Codex, or Claude Code for development, testing, and troubleshooting, with the ability to verify generated results.

Fluent in English and Mandarin.

Preferred Skills

Experience with Ray, Spark, Flink, Kubernetes Jobs, Argo, Slurm, or GPU scheduling platforms.

Experience with cross-region scheduling, GPU utilization optimization, or large-scale asynchronous computing.

Experience with media processing, point cloud processing, simulation workloads, or ML compute platforms.

Equal Opportunity Statement

We are committed to creating a diverse and inclusive environment for all employees.

Apply Now

Similar Opportunities

View all jobs