Epergne Solutions
Singapore / Global
Singapore / Global
Job Title:
DevOps / Site Reliability Engineer (SRE)
Experience:
5?12 Years
Role Overview:
We are looking for an experienced
DevOps / Site Reliability Engineer (SRE)
with strong expertise in
Cloud Infrastructure, Kubernetes, Terraform, and Python scripting . The ideal candidate should have hands-on experience managing scalable and highly available cloud environments, automating infrastructure, and supporting production systems.
Key Responsibilities:
Design, implement, and maintain cloud infrastructure across
AWS, GCP, or Alibaba Cloud
environments.
Build and manage Infrastructure as Code (IaC) using
Terraform .
Develop and maintain automation scripts using
Python .
Deploy, manage, and troubleshoot
Kubernetes
clusters and containerized applications.
Monitor system performance, reliability, and availability, ensuring minimal downtime.
Implement CI/CD pipelines and automate deployment processes.
Collaborate with development and operations teams to improve system scalability and resilience.
Perform incident management, root cause analysis, and production support activities.
Required Skills:
5?12 years of experience in
DevOps / SRE
roles.
Strong hands-on experience with
AWS and/or GCP
cloud platforms.
Experience with
Alibaba Cloud
is an added advantage.
Proficiency in
Python scripting .
Strong experience in
Terraform
for infrastructure automation.
Hands-on experience with
Kubernetes
administration and troubleshooting.
Knowledge of CI/CD tools and DevOps best practices.
Experience in monitoring, logging, and performance tuning of cloud-native applications.
Preferred Skills:
Experience with containerization technologies such as Docker.
Exposure to cloud security and networking concepts.
Strong problem-solving and troubleshooting skills.
#J-18808-Ljbffr
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global