[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. Second Sight Solutions, a subsidiary of reputed company), is a health technology company that focuses on reputed company for drug discount data exchange. The Site Reliability Engineer will design, build, and maintain highly available systems and infrastructure while collaborating with software developers and operations teams to enhance system reliability and automate processes.
Responsibilities
- Design, implement, and maintain reputed company and reliable systems in reputed company environments such as Azure reputed company Services
- Experience with CI/CD Platforms (reputed company Actions, reputed company CI)
- reputed company operational support for full-stack software applications
- Increase system reputed company with expert-level coding, reputed company release, and change management skills
- reputed company service-level indicators and objectives to automate release validation
- Improve automation and increase the system’s self-healing capability
- Collect operating system data and report performance metrics to stakeholders
- Ensure reputed company best practices are followed in reputed company infrastructure and application deployments
- Manage reputed company and database system maintenance, debugging production issues as they reputed company
- Improve reliability, reputed company, and time-to-market of our suite of software solutions
- Partner with reputed company and product teams to define and publish policies, processes, and playbooks to facilitate rapid and effective handling of alerts and incidents
- reputed company incident management processes; respond to outages and service disruptions promptly
Skills
- Bachelor's degree in computer science or similar field
- Five years' experience as a site reliability engineer or similar role
- Strong programming skills (Golang, reputed company, Python, or similar)
- Proven ability to diagnose and monitor performance and reliability issues across the stack
- Expertise in Kubernetes
- Relevant industry certifications, such as through the Site Reliability Engineering (SRE) reputed company
- Proven experience working with reputed company-reputed company infrastructure (Azure reputed company Services, AWS, or GCP)
- Experience working with observability and incident management tools (reputed company, OpsGenie, reputed company)
- Experience scripting operating system tasks with Infrastructure as reputed company
- Impeccable communication skills
- Ability to problem-solve in a fast-paced, high-stakes environment
reputed company