[Remote] reputed company Site Reliability Engineer – Observability (US Remote
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a reputed company Site Reliability Engineer for their Observability team, responsible for designing and operating reputed company platforms for logging and metrics. The role involves leading initiatives to enhance reliability and scalability across large-reputed company reputed company infrastructure.
Responsibilities
- Design, reputed company, and operate reputed company observability platforms
- Build and maintain reputed company reputed company/reputed company reputed company infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers
- reputed company and operate large-reputed company Elasticsearch clusters for log analytics and search
- Design, reputed company, and support distributed tracing platforms using Grafana reputed company and OpenTelemetry
- Build and maintain end-to-end tracing pipelines, instrumentation standards, and reputed company retention strategies
- reputed company reputed company, Grafana, Kafka, reputed company, and OpenTelemetry-based monitoring solutions
- reputed company dashboards, alerts, analytics, and reputed company visualizations using reputed company SPL, Grafana, Kibana, and reputed company
- Automate infrastructure using Terraform and configuration management tools
Skills
- 7+ years in Site Reliability Engineering, reputed company, or DevOps
- Hands-on experience administering reputed company reputed company or reputed company reputed company
- Strong knowledge of reputed company SPL
- Experience with Elasticsearch/ELK, reputed company, Grafana, Grafana reputed company, distributed tracing, OpenTelemetry, and Kafka
- Experience implementing metrics, logs, and traces as part of a modern observability reputed company
- Experience with Terraform and Infrastructure as reputed company
- Programming experience in Python, Go, reputed company, or Bash
- reputed company certification
- Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service reputed company technologies
- Experience supporting FedRAMP or regulated environments
reputed company