Start Your Search Here

Job Search

u3 infotech pte. ltd.

Singapore / Global

Observability Engineer

Job Description

Key Responsibilities:

Observability Strategy and Governance

Define and own Enterprise Observability Architecture aligned with operational resilience mandates (MAS TRM, DORA, APRA CPS 230).

Deploy and optimize observability platforms (Datadog, Dynatrace, Splunk) for full-stack visibility across infra, application, network, and user experience.

Establish governance standards for telemetry data (metrics, logs, traces), ensuring consistency, retention compliance, and security controls.

Integrate observability platforms with incident management, ITSM, and AIOps systems for predictive alerting and anomaly detection.

Reliability Engineering and Automation

Implement SRE frameworks for infrastructure and business-critical applications.

Automate runbooks, alerts, self-healing actions, and auto-remediation workflows via Python, Ansible, and Terraform.

Partner with Application, Infrastructure, and Cyber teams to codify operational reliability into the delivery lifecycle.

Conduct resilience testing, chaos engineering, and capacity validation.

Develop error budget policies and reliability scorecards for key production services.

Cloud Observability and Platform Engineering

Architect and manage observability for Cloud-native workloads in AWS and Azure.

Integrate cloud observability into landing zones and CI/CD pipelines for continuous compliance.

Implement IaC models using Terraform and Ansible for consistent, auditable provisioning.

Collaborate with Cloud, DevOps, and Security teams on real-time telemetry aligned to audit requirements.

Operational Excellence and Stakeholder Management

Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.

Deliver executive dashboards highlighting availability, reliability KPIs, and operational risk indicators.

Act as technical advisor to senior management during major incidents, post-incident reviews, and audits.

Skillset Requirements:

At least 5 years of experience in Infrastructure, Cloud, or Site Reliability Engineering (SRE) related roles, with minimum 3 years in an SRE SME capacity, ideally within financial institutions or regulated environments.

Hands-on expertise with Observability Platforms: Datadog, Dynatrace, Splunk, ELK.

Hands-on expertise with Automation/IaC: Terraform, Ansible, Python, CI/CD tools.

Hands-on expertise with Cloud Platforms: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).

Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.

Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.

A good team player with excellent written and verbal communication skills, able to coordinate across diverse stakeholders.

Certification in at least one of the following required:

Datadog Certified Observability Professional / Dynatrace Certified Associate

Terraform/Ansible/Python Certified Expert

Certification in the following will be advantageous:

AWS Certified DevOps Engineer / Azure DevOps Expert

SRE Foundation/Practitioner (DevOps Institute)

ITIL v4 Managing Professional

U3 Privacy Notice:

Please refer to U3's Privacy Notice for Job Applicants/Seekers at When you apply, you voluntarily consent to the collection, use and disclosure of your personal data for recruitment/employment and related purposes.

Apply Now

Similar Opportunities

View all jobs