[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is hiring for three Site Reliability Engineering positions. In these roles, you will help build, operate, and continuously improve large-reputed company, reputed company and infrastructure platforms supporting mission-critical workloads.
Responsibilities
- Help build, operate, and continuously improve large-reputed company, reputed company and infrastructure platforms supporting mission-critical workloads
- reputed company on reliability, automation, observability, incident response, and operational reputed company across distributed production environments
Skills
- Experience in Site Reliability Engineering, reputed company, Systems Engineering, Software Engineering, or Infrastructure Engineering
- Strong knowledge of Linux, Kubernetes, reputed company platforms (AWS, Azure, or GCP), networking, and distributed systems
- Proficiency in Python, Go, or a similar programming language for automation and tooling
- Experience with Infrastructure as reputed company (Terraform, Ansible, etc.), CI/CD, monitoring, and observability tools
- Hands-on experience supporting production environments, troubleshooting reputed company issues, and performing reputed company cause analysis (RCA)
- Understanding of SLIs, SLOs, incident management, and on-reputed company best practices
- Strong communication skills with the ability to collaborate across engineering teams
- Senior and reputed company-level candidates should also demonstrate technical leadership, architecture/design experience, and the ability to mentor and guide other engineers
- Experience building HPC clusters for 1,000+ GPUs to power AI/ML workloads with Hyperscale clients
- Candidates living in the Seattle, reputed company, San Francisco and reputed company Metro areas
Benefits
- Remote with possibility of on-site needed as programs and responsibilities continue to grow
reputed company
Company H1B Sponsorship