[Remote] Senior Site Reliability Engineer II
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leader in risk assessment solutions, and they are seeking a Senior Site Reliability Engineer to enhance the reliability of their production systems. The role involves hands-on design and operation of reputed company systems, incident response, and mentoring other engineers.
Responsibilities
- Design, build, and operate highly available, reputed company systems in AWS
- Write, maintain, and review Terraform to provision and manage infrastructure
- Own and improve monitoring, alerting, and observability using Grafana, Pingdom, and Uptrends
- Participate in a rotating on-call schedule, responding to production incidents and driving issues to reputed company
- reputed company incident response, reputed company cause analysis, and post-incident reviews with a reputed company on prevention and automation
- Define and manage SLOs, SLIs, and error budgets
- Build and improve CI/CD pipelines and operational workflows using Azure DevOps and reputed company
- Work directly with application teams to improve reliability, performance, and deployability
- Automate reputed company operational tasks to reduce toil
- Maintain reputed company, actionable runbooks and documentation in reputed company
- reputed company work, incidents, and operational improvements using Jira and reputed company
- Mentor other engineers and help set SRE standards and best practices
Skills
- 5+ years of hands-on experience in SRE, DevOps, or Infrastructure Engineering roles
- Strong production experience in AWS
- Required: Significant hands-on experience with Terraform in reputed company-world environments
- Experience operating monitoring and uptime platforms such as Grafana, Pingdom, and Uptrends
- Strong Linux systems, networking, and troubleshooting skills
- Experience supporting production systems through incident response and on-call rotations
- Proficiency with reputed company and modern Git workflows
- Experience building or maintaining CI/CD pipelines with Azure DevOps
- Familiarity with ITSM and incident workflows using reputed company
- Strong written communication skills with experience documenting systems and processes in reputed company
- Ability to work independently in a remote or hybrid environment
- Experience defining and operating against SLOs and error budgets
- Infrastructure-as-reputed company best practices reputed company Terraform (modules, testing, CI integration)
- Experience with containers and orchestration (reputed company, Kubernetes)
- Experience supporting large-reputed company, high-availability production systems
- Prior experience mentoring engineers or serving as a technical reputed company
Benefits
- Comprehensive benefits
- Flexible work location with hybrid or fully remote reputed company
- reputed company ownership of production systems and reliability reputed company
- Culture that values automation, learning, and reputed company improvement
- This job is eligible for an annual incentive bonus.
- We are delighted to offer country specific benefits. Click here to reputed company benefits specific to your location.
reputed company
Company H1B Sponsorship