Monitoring and Observability Specialist

Yerevan, Armenia · Full time

We are looking for a Monitoring and Observability Specialist to design, build, and maintain end-to-end observability infrastructure across our platform. In this role, you will own the full lifecycle of monitoring, logging, and tracing systems — from architecture and implementation to custom agent development and ongoing optimization. You will work closely with engineering, DevOps, and platform teams to ensure deep visibility into every layer of our infrastructure and applications.

What we offer
  • An open, feedback-driven culture that supports continuous professional growth, whether as a technical leader or people manager.
  • A distributed, hybrid work environment with teams across multiple time zones.
  • Competitive compensation and well-being benefits.
  • The opportunity to shape observability strategy from the ground up and pilot modern tooling and practices.
  • A collaborative environment where your expertise directly impacts platform reliability and performance.
What you will do
  • Design and architect end-to-end observability solutions covering metrics, logs, traces, and synthetic monitoring across on-premise and cloud environments.
  • Build, maintain, and extend custom monitoring agents and exporters tailored to internal services and infrastructure components.
  • Develop and maintain high-fidelity dashboards, alerts, and runbooks that provide actionable insights to both operations and engineering teams.
  • Own the full lifecycle of the monitoring stack — deployment, scaling, performance tuning, and capacity planning.
  • Implement and maintain distributed tracing pipelines to enable deep visibility into service-to-service interactions.
  • Define and enforce observability standards and best practices across teams.
  • Participate in incident response, leveraging observability data to accelerate root-cause analysis and resolution.
  • Evaluate and integrate new monitoring and observability tooling as the ecosystem evolves.
  • Collaborate closely with platform, networking, and application teams to ensure comprehensive coverage and minimal blind spots.
About you
  • You are a hands-on engineer with deep expertise in building and operating observability platforms at scale.
  • You have a strong understanding of Linux internals and networking, and you are comfortable writing production-grade code.
  • You thrive in complex, multi-layered infrastructure environments and can translate raw telemetry data into meaningful operational intelligence.
Must-have
  • 5+ years of experience in monitoring, observability, or SRE roles with hands-on ownership of observability infrastructure.
  • Deep expertise in at least two of: Prometheus / VictoriaMetrics, Elastic Stack (Elasticsearch, Logstash, Kibana, Filebeat), Grafana, Loki, Jaeger / Tempo / OpenTelemetry.
  • Strong experience writing and maintaining custom monitoring agents or exporters (Go, Python, or similar).
  • Solid Linux administration and internal knowledge (networking stack, kernel tuning, process management).
  • Strong networking fundamentals — TCP/IP, DNS, VLANs, routing, traffic analysis, firewall rules.
  • Experience deploying and operating monitoring systems in Kubernetes environments.
  • Proficiency in IaC and GitOps tooling (Terraform, Ansible, Flux, ArgoCD, or similar).
  • Scripting and automation skills (Python, Go, Bash). Experience with log aggregation pipelines and understanding of log management at scale.
  • Strong incident response experience and ability to build alerting strategies that reduce noise and improve MTTR.
Nice to have
  • Experience with distributed tracing systems (Jaeger, Tempo, OpenTelemetry) at scale.
  • Familiarity with time-series database internals and query optimization (PromQL, MetricsQL, KQL).
  • Experience with service mesh observability (Istio, Linkerd telemetry).
  • Knowledge of APM and synthetic monitoring approaches.
  • Familiarity with streaming platforms (Kafka) as part of log/metrics pipelines.
  • Experience with cloud-native monitoring in multi-cloud or hybrid environments (AWS, GCP, Azure).
  • Background in capacity planning and performance engineering.