Start Your Search Here

Job Search

websparks pte. ltd.

Singapore / Global

Observability Engineer (Public Sector)

  • Hybrid

Job Description

1-year contract, renewable

Government project

Hybrid work arrangement

We are looking for an Observability Engineer to help build and operate the observability capabilities of the future SSOE platform. You will help provide end-to-end visibility across MOE's technology environment, spanning on-premises infrastructure, networks, applications, cloud platforms, and hybrid environments. You will enable engineering and operations teams to understand system health, identify issues early, diagnose incidents quickly, and continuously improve service reliability.

What You Will Be Working On

As an Observability Engineer, you will establish and operate consistent observability capabilities across SSOE infrastructure and applications.

You will work across metrics, events, logs, and traces to provide a unified view of service health and performance. You will define instrumentation standards, service-level indicators and objectives, alerting strategies, dashboards, and operational health signals.

You will work closely with application, infrastructure, network, security, and platform teams to ensure observability is built into services from the outset rather than added after deployment.

Key Responsibilities

Observability Engineering

Design and operate end-to-end observability across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments

Collect, aggregate, and correlate metrics, events, logs, and traces across infrastructure and application workloads

Define and maintain observability standards that work consistently across legacy, on-premise, containerised, and cloud-native systems

Establish application and infrastructure instrumentation standards using OpenTelemetry and other appropriate technologies

Support engineering teams with instrumentation, SDK, agent, and telemetry integration

Define common conventions for service naming, metadata, tagging, correlation IDs, and telemetry enrichment

Identify observability gaps and continuously improve end-to-end visibility across SSOE services

Service Reliability & Monitoring

Define SLIs, SLOs, alerting rules, and service health indicators for critical services

Build operational dashboards covering infrastructure health, application performance, user experience, availability, capacity, and service reliability

Develop leadership-level views that provide meaningful visibility into service performance and operational trends

Design actionable alerting that enables teams to identify and respond to issues while minimising unnecessary alert noise

Establish monitoring and operational-readiness requirements for new applications, infrastructure, and platform components

Use observability data to support capacity planning, performance analysis, reliability improvements, and operational decision-making

Telemetry & Integration

Define secure telemetry collection and routing across on-premise environments, GCC, cloud platforms, and approved SaaS services

Work with infrastructure and platform teams to integrate telemetry from servers, network devices, applications, containers, databases, and managed cloud services

Design observability approaches that account for network boundaries, security zones, data residency, and connectivity constraints

Define telemetry retention, lifecycle, and cost-management requirements

Ensure logs, metrics, and traces can be correlated across distributed and hybrid systems

Incident Management & Continuous Improvement

Support operational teams during incidents by using observability data to identify symptoms, dependencies, and potential root causes

Participate in incident investigation, root-cause analysis, and post-incident reviews

Identify recurring operational issues and recommend improvements to instrumentation, alerting, architecture, or operational processes

Define appropriate SLOs and operational health indicators for observability services

Participate in operational support and on-call responsibilities for owned services

Maintain architecture documentation, operational procedures, and runbooks

What We Are Looking For

Experience

Minimum

3-5 years of experience

in observability engineering, Site Reliability Engineering (SRE), platform engineering, infrastructure engineering, or a related discipline

At least

2 years of hands-on experience

implementing or operating observability and monitoring capabilities in production environments

Demonstrated experience working with

metrics, logging, tracing, dashboards, alerting, and incident troubleshooting

Experience monitoring and supporting

production infrastructure, applications, or distributed systems

Experience working with

on-premise and/or cloud environments , with an understanding of hybrid infrastructure

Experience working with engineering or operations teams to implement instrumentation and improve service reliability

Technical Skills

Observability: Dynatrace, Elastic, Grafana, Prometheus or equivalent

Instrumentation: OpenTelemetry

Cloud: AWS-native observability and monitoring capabilities

Infrastructure: Docker, ECS, CI/CD, SHIP-HATS, Terraform / OpenTofu, Ansible

Data & Storage: Cloud-native storage and telemetry data services

Event & Telemetry: Kafka, MQ, event-driven telemetry patterns

Engineering Practices

Treats observability configuration, instrumentation, dashboards, and platform components as version-controlled engineering artefacts

Uses automation and Infrastructure as Code for repeatable and auditable deployments

Designs observability for reliability, scalability, security, and operational sustainability

Understands the difference between collecting telemetry and creating useful operational signals

Designs monitoring and alerting around service and user impact rather than individual infrastructure metrics alone

Builds observability into services from the beginning of the engineering lifecycle

Uses code review, testing, and CI/CD for observability-related changes where appropriate

Collaborates effectively across application, infrastructure, network, security, data, and platform teams

Security & Governance

Apply MOE and Government data-classification requirements

Ensure telemetry is collected, transmitted, stored, and accessed according to applicable security and data-residency requirements

Prevent sensitive information, credentials, and secrets from being unnecessarily captured in telemetry

Implement appropriate access controls for observability platforms and operational data

Maintain auditability and traceability of observability configuration and operational activities

Participate in security, architecture, and operational-readiness reviews

Nice to Have

Experience with Singapore Government platforms such as TechPass, SHIP-HATS, SEED, and GCC

Familiarity with OC/SN data-classification requirements

AWS or Azure cloud certifications

Experience implementing OpenTelemetry at scale

Experience operating observability platforms across large or distributed environments

Experience monitoring hybrid infrastructure spanning data centres and cloud environments

Familiarity with SRE practices such as error budgets, SLO management, incident response, and reliability engineering

Apply Now

Similar Opportunities

View all jobs