The Observability Engineer will play a key role in building and continuously improving the organisation’s observability capabilities, helping teams gain greater visibility into the health, performance, and reliability of their systems.
Working across SRE, Platform Engineering, and Application Development, you will help establish effective monitoring, logging, metrics, tracing, and alerting practices that enable teams to identify issues early, troubleshoot efficiently, and improve overall system reliability.
This is a global role with a strong focus on automation, continuous improvement, and creating scalable observability solutions across platforms and services.
Your daily tasks and responsibilities are the following:
- Design, implement, maintain, and evolve enterprise observability platforms.
- Implement and support logging, metrics, tracing, monitoring, and alerting solutions.
- Contribute to the implementation, rollout, maintenance, and continuous improvement of Elastic / ELK and other observability platforms.
- Define instrumentation standards for applications, infrastructure, and cloud services.
- Build dashboards aligned with SLOs and service health indicators.
- Optimise alerting frameworks to improve signal quality and reduce unnecessary alert noise.
- Ensure telemetry pipelines are scalable, reliable, secure, and cost-efficient.
- Monitor the health and reliability of the observability platform itself.
- Analyse telemetry data to identify performance trends, recurring issues, and opportunities for improvement.
- Partner with SRE teams to improve visibility into error budgets, performance trends, and reliability.
- Collaborate with Application Development teams to embed observability into application and solution design.
- Support incident investigation through effective monitoring, dashboards, alerting, and diagnostic capabilities.
- Develop automation and monitoring solutions using technologies such as Python and Bash.
- Work with Infrastructure and Platform teams to integrate observability across cloud-native and distributed environments.
- Contribute to observability standards, best practices, and technical documentation.
- Continuously improve detection and diagnostic capabilities to enable faster identification and resolution of issues.
- Work collaboratively with globally distributed teams across different regions and time zones.