websparks pte. ltd.
Singapore / Global
Observability Engineer (Public Sector)
- Hybrid
Singapore / Global
1-year contract, renewable
Government project
Hybrid work arrangement
We are looking for an Observability Engineer to help build and operate the observability capabilities of the future SSOE platform. You will help provide end-to-end visibility across MOE's technology environment, spanning on-premises infrastructure, networks, applications, cloud platforms, and hybrid environments. You will enable engineering and operations teams to understand system health, identify issues early, diagnose incidents quickly, and continuously improve service reliability.
What You Will Be Working On
As an Observability Engineer, you will establish and operate consistent observability capabilities across SSOE infrastructure and applications.
You will work across metrics, events, logs, and traces to provide a unified view of service health and performance. You will define instrumentation standards, service-level indicators and objectives, alerting strategies, dashboards, and operational health signals.
You will work closely with application, infrastructure, network, security, and platform teams to ensure observability is built into services from the outset rather than added after deployment.
Key Responsibilities
Observability Engineering
Design and operate end-to-end observability across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments
Collect, aggregate, and correlate metrics, events, logs, and traces across infrastructure and application workloads
Define and maintain observability standards that work consistently across legacy, on-premise, containerised, and cloud-native systems
Establish application and infrastructure instrumentation standards using OpenTelemetry and other appropriate technologies
Support engineering teams with instrumentation, SDK, agent, and telemetry integration
Define common conventions for service naming, metadata, tagging, correlation IDs, and telemetry enrichment
Identify observability gaps and continuously improve end-to-end visibility across SSOE services
Service Reliability & Monitoring
Define SLIs, SLOs, alerting rules, and service health indicators for critical services
Build operational dashboards covering infrastructure health, application performance, user experience, availability, capacity, and service reliability
Develop leadership-level views that provide meaningful visibility into service performance and operational trends
Design actionable alerting that enables teams to identify and respond to issues while minimising unnecessary alert noise
Establish monitoring and operational-readiness requirements for new applications, infrastructure, and platform components
Use observability data to support capacity planning, performance analysis, reliability improvements, and operational decision-making
Telemetry & Integration
Define secure telemetry collection and routing across on-premise environments, GCC, cloud platforms, and approved SaaS services
Work with infrastructure and platform teams to integrate telemetry from servers, network devices, applications, containers, databases, and managed cloud services
Design observability approaches that account for network boundaries, security zones, data residency, and connectivity constraints
Define telemetry retention, lifecycle, and cost-management requirements
Ensure logs, metrics, and traces can be correlated across distributed and hybrid systems
Incident Management & Continuous Improvement
Support operational teams during incidents by using observability data to identify symptoms, dependencies, and potential root causes
Participate in incident investigation, root-cause analysis, and post-incident reviews
Identify recurring operational issues and recommend improvements to instrumentation, alerting, architecture, or operational processes
Define appropriate SLOs and operational health indicators for observability services
Participate in operational support and on-call responsibilities for owned services
Maintain architecture documentation, operational procedures, and runbooks
What We Are Looking For
Experience
Minimum
3-5 years of experience
in observability engineering, Site Reliability Engineering (SRE), platform engineering, infrastructure engineering, or a related discipline
At least
2 years of hands-on experience
implementing or operating observability and monitoring capabilities in production environments
Demonstrated experience working with
metrics, logging, tracing, dashboards, alerting, and incident troubleshooting
Experience monitoring and supporting
production infrastructure, applications, or distributed systems
Experience working with
on-premise and/or cloud environments , with an understanding of hybrid infrastructure
Experience working with engineering or operations teams to implement instrumentation and improve service reliability
Technical Skills
Observability: Dynatrace, Elastic, Grafana, Prometheus or equivalent
Instrumentation: OpenTelemetry
Cloud: AWS-native observability and monitoring capabilities
Infrastructure: Docker, ECS, CI/CD, SHIP-HATS, Terraform / OpenTofu, Ansible
Data & Storage: Cloud-native storage and telemetry data services
Event & Telemetry: Kafka, MQ, event-driven telemetry patterns
Engineering Practices
Treats observability configuration, instrumentation, dashboards, and platform components as version-controlled engineering artefacts
Uses automation and Infrastructure as Code for repeatable and auditable deployments
Designs observability for reliability, scalability, security, and operational sustainability
Understands the difference between collecting telemetry and creating useful operational signals
Designs monitoring and alerting around service and user impact rather than individual infrastructure metrics alone
Builds observability into services from the beginning of the engineering lifecycle
Uses code review, testing, and CI/CD for observability-related changes where appropriate
Collaborates effectively across application, infrastructure, network, security, data, and platform teams
Security & Governance
Apply MOE and Government data-classification requirements
Ensure telemetry is collected, transmitted, stored, and accessed according to applicable security and data-residency requirements
Prevent sensitive information, credentials, and secrets from being unnecessarily captured in telemetry
Implement appropriate access controls for observability platforms and operational data
Maintain auditability and traceability of observability configuration and operational activities
Participate in security, architecture, and operational-readiness reviews
Nice to Have
Experience with Singapore Government platforms such as TechPass, SHIP-HATS, SEED, and GCC
Familiarity with OC/SN data-classification requirements
AWS or Azure cloud certifications
Experience implementing OpenTelemetry at scale
Experience operating observability platforms across large or distributed environments
Experience monitoring hybrid infrastructure spanning data centres and cloud environments
Familiarity with SRE practices such as error budgets, SLO management, incident response, and reliability engineering
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global
Singapore / Global