u3 infotech pte. ltd.
Singapore / Global
Singapore / Global
Key Responsibilities:
Observability Strategy and Governance
Define and own Enterprise Observability Architecture aligned with operational resilience mandates (MAS TRM, DORA, APRA CPS 230).
Deploy and optimize observability platforms (Datadog, Dynatrace, Splunk) for full-stack visibility across infra, application, network, and user experience.
Establish governance standards for telemetry data (metrics, logs, traces), ensuring consistency, retention compliance, and security controls.
Integrate observability platforms with incident management, ITSM, and AIOps systems for predictive alerting and anomaly detection.
Reliability Engineering and Automation
Implement SRE frameworks for infrastructure and business-critical applications.
Automate runbooks, alerts, self-healing actions, and auto-remediation workflows via Python, Ansible, and Terraform.
Partner with Application, Infrastructure, and Cyber teams to codify operational reliability into the delivery lifecycle.
Conduct resilience testing, chaos engineering, and capacity validation.
Develop error budget policies and reliability scorecards for key production services.
Cloud Observability and Platform Engineering
Architect and manage observability for Cloud-native workloads in AWS and Azure.
Integrate cloud observability into landing zones and CI/CD pipelines for continuous compliance.
Implement IaC models using Terraform and Ansible for consistent, auditable provisioning.
Collaborate with Cloud, DevOps, and Security teams on real-time telemetry aligned to audit requirements.
Operational Excellence and Stakeholder Management
Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.
Deliver executive dashboards highlighting availability, reliability KPIs, and operational risk indicators.
Act as technical advisor to senior management during major incidents, post-incident reviews, and audits.
Skillset Requirements:
At least 5 years of experience in Infrastructure, Cloud, or Site Reliability Engineering (SRE) related roles, with minimum 3 years in an SRE SME capacity, ideally within financial institutions or regulated environments.
Hands-on expertise with Observability Platforms: Datadog, Dynatrace, Splunk, ELK.
Hands-on expertise with Automation/IaC: Terraform, Ansible, Python, CI/CD tools.
Hands-on expertise with Cloud Platforms: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).
Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.
Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
A good team player with excellent written and verbal communication skills, able to coordinate across diverse stakeholders.
Certification in at least one of the following required:
Datadog Certified Observability Professional / Dynatrace Certified Associate
Terraform/Ansible/Python Certified Expert
Certification in the following will be advantageous:
AWS Certified DevOps Engineer / Azure DevOps Expert
SRE Foundation/Practitioner (DevOps Institute)
ITIL v4 Managing Professional
U3 Privacy Notice:
Please refer to U3's Privacy Notice for Job Applicants/Seekers at When you apply, you voluntarily consent to the collection, use and disclosure of your personal data for recruitment/employment and related purposes.
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global