Observability Engineer

  • Permanent

COCUS ProSource

COCUS Prosource is all about People! We are proud to deliver skilled services and products developed by great talent, with attitude and ambition to work in innovative IT solutions.

Emotions are part of us, we encourage everyone to be what they truly are in a collaborative, informal, transparent, and open environment, that is why we take our partnerships seriously - supporting as a Talent Acquisition specialized partner on the recruitment for companies with the same People first mindset as we have!

What you will be doing

The Observability Engineer will play a key role in building and continuously improving the organisation’s observability capabilities, helping teams gain greater visibility into the health, performance, and reliability of their systems.

Working across SRE, Platform Engineering, and Application Development, you will help establish effective monitoring, logging, metrics, tracing, and alerting practices that enable teams to identify issues early, troubleshoot efficiently, and improve overall system reliability.

This is a global role with a strong focus on automation, continuous improvement, and creating scalable observability solutions across platforms and services.

Your daily tasks and responsibilities are the following:

  • Design, implement, maintain, and evolve enterprise observability platforms.
  • Implement and support logging, metrics, tracing, monitoring, and alerting solutions.
  • Contribute to the implementation, rollout, maintenance, and continuous improvement of Elastic / ELK and other observability platforms.
  • Define instrumentation standards for applications, infrastructure, and cloud services.
  • Build dashboards aligned with SLOs and service health indicators.
  • Optimise alerting frameworks to improve signal quality and reduce unnecessary alert noise.
  • Ensure telemetry pipelines are scalable, reliable, secure, and cost-efficient.
  • Monitor the health and reliability of the observability platform itself.
  • Analyse telemetry data to identify performance trends, recurring issues, and opportunities for improvement.
  • Partner with SRE teams to improve visibility into error budgets, performance trends, and reliability.
  • Collaborate with Application Development teams to embed observability into application and solution design.
  • Support incident investigation through effective monitoring, dashboards, alerting, and diagnostic capabilities.
  • Develop automation and monitoring solutions using technologies such as Python and Bash.
  • Work with Infrastructure and Platform teams to integrate observability across cloud-native and distributed environments.
  • Contribute to observability standards, best practices, and technical documentation.
  • Continuously improve detection and diagnostic capabilities to enable faster identification and resolution of issues.
  • Work collaboratively with globally distributed teams across different regions and time zones.

What we are looking for

  • Several years of experience in Monitoring, Observability, SRE, Platform Engineering, or related roles.   
  • Strong hands-on experience with observability platforms (e.g., Azure Monitor, Prometheus, Datadog, Splunk, ideally Elastic / ELK).
  • Strong experience with Azure and cloud-native monitoring integrations.
  • Experience implementing and supporting enterprise-level logging, metrics, and monitoring platforms.
  • Experience designing or supporting scalable telemetry architectures and pipelines.
  •  Strong knowledge of alert design and signal-to-noise optimisation.
  • Understanding of SLOs, SLIs, error budgets, and SLO-driven monitoring strategies.
  • Experience working with cloud-native and distributed systems environments.
  • Experience with Terraform, Ansible, GitHub Actions, Azure DevOps, or similar DevOps and automation tools.
  • Scripting or automation capabilities using Python, Bash, or similar languages.
  •  Strong analytical, problem-solving, communication, and collaboration skills, with the ability to work effectively across technical teams.
  •  Comfortable working in a global environment, collaborating across different cultures, geographies, and time zones.
  • Able to thrive in a changing environment, take ownership, and deal effectively with ambiguity.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • Fluent in written and spoken English.

What this opportunity offers

  • Permanent contract with long-term career prospects 
  • Competitive salary aligned with your experience and expertise 
  • Annual target bonus based on performance 
  • Daily meal allowance 
  • Health insurance and Life insurance 
  • Access to Employee Assistance Program for well-being and support 
  • A flexible hybrid work model