optimum solutions (singapore) pte ltd
Singapore / Global
Singapore / Global
Job Description:
Essential Key Technical Requirement:
Role: Datadog Engineer
Possess a degree in Computer Science/Information Technology or related fields.
At least 5 years of experience in Infrastructure, Cloud, or Site Reliability Engineering (SRE) related roles, with minimum 3 years in an SRE SME capacity, ideally within financial institutions or regulated environments.
Mandatory is
: Should have recent 2+ years hands-on working experience asa Datadog administrator who can manage the Datadog platformend-to-end
Must have recent working experience in Datadog platform end-to-end
Must: Hands-on expertise with Observability Platforms: Datadog, Dynatrace, Splunk, ELK.
Must: Hands-on expertise with Automation/IaC: Terraform, Ansible, Python, CI/CD tools.
Must: Hands-on expertise with Cloud Platforms: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).
Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.
Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
Certification in at least
one of the following required
:
Datadog CertifiedObservability Professional / Dynatrace Certified Associate
Terraform/Ansible/Python Certified Expert
Certification in the
following will be advantageous:
AWSCertified DevOps Engineer / Azure DevOps Expert
SRE Foundation/Practitioner (DevOps Institute)ITIL v4 Managing Professional
Project Summary:
Develop and execute enterprise observability and service reliability strategy across all infrastructure and application domains, driving proactive monitoring, automation, and resilience initiatives.
Key Responsibilities:
ObservabilityStrategy and Governance:
Define and own Enterprise Observability Architecture aligned with operational resilience mandates (MAS TRM, DORA, APRA CPS 230).
Deploy and optimize observability platforms (Datadog, Dynatrace, Splunk) for full-stack visibility across infra, application, network, and user experience.
ReliabilityEngineering and Automation:
Implement SRE frameworks for infrastructure and business-critical applications.
Automate runbooks, alerts, self-healing actions, and auto-remediationworkflows via Python, Ansible, and Terraform.
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global