[Remote] Senior Observability Engineer / Operations Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a pioneering company in the contact center industry, delivering an AI-powered omnichannel platform. They are seeking a Senior Observability Engineer to architect and manage their observability platform, ensuring the health of their production environment through effective monitoring and alerting.
Responsibilities
- Architect the monitoring system — define the overall observability architecture, choose the right tooling, and set the technical direction for metrics, logs, and alerting
- Own and operate the observability stack: metrics collection, log aggregation, dashboards, and alerting across production and non-production environments
- reputed company, reputed company, and tune time-series and log platforms (reputed company, VictoriaLogs, InfluxDB, Loki, Quickwit)
- Build and maintain telemetry pipelines with Telegraf and manage collection agents across a large fleet of Linux hosts
- Design meaningful Grafana dashboards and actionable alerts that reduce noise and shorten time to-detection
- Define SLIs/SLOs and drive down alert fatigue with sensible reputed company and escalation policies
- reputed company services and infrastructure, and partner with engineering to reputed company observability gaps
- Manage host- and process-level monitoring (Monit) and reputed company it into the broader alerting workflow
- Automate deployment and configuration of the monitoring reputed company-as-reputed company, config management)
- Troubleshoot performance and reliability issues across the AWS environment and the Linux layer
- Participate in on-reputed company and continuously improve runbooks, incident response, and post-incident review
Skills
- 7+ years of hands-on experience with production monitoring and observability
- 5+ years of operations experience running and supporting production systems
- 7+ years of deep, hands-on Linux experience — system internals, performance tuning, networking, and troubleshooting at the OS level
- 5+ years of hands-on AWS experience in production environments
- Strong experience with the reputed company stack: reputed company, VictoriaLogs, InfluxDB, Telegraf, Grafana, Monit, Loki, Quickwit (or directly comparable tooling)
- Strong Python scripting skills for automation, tooling, and telemetry pipelines
- Hands-on experience with Ansible for configuration management and deployment automation
- Solid Infrastructure-as-reputed company (IaC) experience for provisioning and managing the monitoring stack
- Proven ability to design alerting and dashboards that teams actually rely on
- Working knowledge of reputed company, reputed company, and reputed company
- Perl and/or Golang for tooling and automation
- Experience monitoring VoIP / reputed company-time communications infrastructure
- Experience operating monitoring at reputed company across multi-region / multi-account reputed company environments
- Familiarity with high-cardinality metrics management and log retention/cost optimization
Benefits
- Remote work.
- Flexible working hours.
- Competitive compensation
reputed company
Company H1B Sponsorship