About the role
Where you will make an impact
You will help customers understand what their systems are doing and make production failures easier to detect, explain, and prevent. The work spans OpenTelemetry, metrics, logs, traces, dashboards, alerting, SLOs, incident analysis, Kubernetes, cloud monitoring, and the practical engineering needed to keep telemetry reliable.
What you will do
Core responsibilities
- Implement and operate OpenTelemetry instrumentation and Collector pipelines for metrics, logs, and traces.
- Build and improve dashboards, alerts, service-level indicators, service-level objectives, and incident diagnostics.
- Work with Prometheus, Grafana, cloud-native monitoring services, and commercial observability platforms where appropriate for the customer.
- Investigate missing telemetry, noisy alerts, recurring incidents, latency, capacity, and reliability problems to root cause.
- Support Kubernetes and cloud workload observability and automate repeatable reliability tasks with scripting and infrastructure as code.
- Document telemetry architecture, operational runbooks, service ownership, and customer recommendations clearly.
What we are looking for
Role requirements
- Practical production experience in SRE, DevOps, cloud operations, platform engineering, monitoring, or observability, typically two or more years.
- Working knowledge of Linux, networking fundamentals, Docker, Kubernetes, APIs, Git, and at least one scripting language such as Python or Bash.
- Hands-on understanding of metrics, logs, traces, alerting, incident troubleshooting, Prometheus, Grafana, and OpenTelemetry concepts.
- Ability to follow a telemetry path from application instrumentation through collectors and exporters into a backend and troubleshoot failures systematically.
- A degree in Computer Science, Engineering, Information Technology, or a related discipline is welcome but not required where equivalent professional experience is demonstrated.
- CKA, AWS or Azure certifications, Terraform Associate, Grafana or observability credentials are preferred rather than mandatory.
Selection process
Practical technical assessment
For technical roles, Onyx evaluates practical engineering judgement alongside experience and certifications. Shortlisted candidates may be asked to complete or discuss work such as:
- Instrument or review a small application and explain how telemetry should move through an OpenTelemetry Collector into an observability backend.
- Troubleshoot a deliberately broken telemetry path, such as a Collector, exporter, endpoint, context-propagation, or Kubernetes connectivity problem.
- Create or explain a useful dashboard, alert, SLI, and SLO and justify why each signal matters to the service owner.
What success looks like
Outcomes that matter
- Critical services expose useful telemetry across metrics, logs, and traces.
- Alerts become more actionable and service-level objectives reflect meaningful customer or business outcomes.
- Missing telemetry and recurring incidents are investigated systematically to root cause.
- Customer teams receive clear runbooks, dashboards, and reliability recommendations they can operate after handover.
Development at Onyx
Where the role can grow
- Grow toward Senior SRE, Observability Lead, or Reliability Engineering leadership.
- Develop multi-platform depth across OpenTelemetry and the observability tools used by Onyx customers.
- Pursue Kubernetes, cloud, Terraform, SRE, and observability certifications with Onyx-supported development.
Why Onyx
Clear standards. Real ownership.
Onyx Technologies Group is building from Ghana across technology, software, cloud, financial technology, and emerging businesses. You will work with teams that value practical execution, technical quality, customer outcomes, and the responsibility to build systems that can scale.
Start a conversation