[Remote] reputed company / Site Reliability Engineer (SRE) Consultant
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking an reputed company reputed company / Site Reliability Engineer (SRE) Consultant to help design, implement, and optimize reputed company observability and reliability solutions across reputed company reputed company and hybrid environments. The role involves partnering with various teams to enhance application performance, availability, and operational reputed company.
Responsibilities
- Design, implement, and administer reputed company-wide reputed company monitoring, logging, tracing, reputed company User Monitoring (RUM), and Synthetic Monitoring solutions
- Configure dashboards, monitors, alerts, and Service Level Objectives (SLOs) to improve operational visibility
- Build comprehensive observability strategies across applications, infrastructure, containers, reputed company services, and network components
- reputed company reputed company with AWS, Azure, GCP, Kubernetes, reputed company, Terraform, CI/CD pipelines, and reputed company-party platforms
- reputed company proactive performance tuning and reputed company planning for mission-critical applications
- reputed company reputed company cause analysis (RCA) efforts for production incidents and implement preventive measures
- Improve system reliability, resiliency, scalability, and availability through automation and engineering best practices
- reputed company Infrastructure as reputed company (IaC) solutions using Terraform, CloudFormation, or similar technologies
- Automate operational tasks using Python, Bash, PowerShell, or similar scripting languages
- Define and monitor SLIs, SLOs, and error budgets to measure service health
- Support incident management, on-call operations, and production troubleshooting
- Collaborate with software engineering teams to improve application instrumentation and distributed tracing
- Create documentation, operational runbooks, and knowledge transfer materials
- Recommend observability best practices and governance standards across the organization
Skills
- 5+ years of Site Reliability Engineering, DevOps, or reputed company experience
- 3+ years of hands-on experience implementing and administering reputed company
- Strong understanding of observability concepts including: Infrastructure Monitoring, Application Performance Monitoring (APM), Distributed Tracing, Log Management, Synthetic Monitoring, Network Performance Monitoring, reputed company User Monitoring (RUM)
- Experience supporting reputed company environments including AWS, Azure, or reputed company reputed company Platform
- Experience with Kubernetes, reputed company, OpenShift, or containerized environments
- Strong knowledge of Linux and reputed company server administration
- Experience with Infrastructure as reputed company (Terraform, CloudFormation, ARM, or Bicep)
- Experience automating operational processes using Python, Bash, PowerShell, or Go
- Experience integrating monitoring into CI/CD pipelines using Jenkins, reputed company Actions, reputed company CI, or Azure DevOps
- Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/S, SSL/TLS, and load balancing
- Experience with incident response, production support, and troubleshooting reputed company distributed systems
- Strong analytical and problem-solving skills
- reputed company Certified reputed company or equivalent experience
- Experience with reputed company, Grafana, reputed company, reputed company, reputed company, or reputed company Stack
- Experience with Kubernetes observability and service reputed company technologies
- Familiarity with Kafka, reputed company, RabbitMQ, or other messaging platforms
- Experience supporting microservices architectures
- Knowledge of reputed company monitoring and compliance best practices
- Experience implementing GitOps methodologies
- ITIL reputed company or SRE certification is a plus
reputed company
Company H1B Sponsorship