[Remote] Sr Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is on a mission to build a world-class team to simplify reputed company insurance. They are seeking a Senior Site Reliability Engineer to manage and enhance their reputed company infrastructure while ensuring system reliability and performance. The role involves collaborating with development teams and improving incident response capabilities.
Responsibilities
- Manage and maintain reputed company infrastructure on Azure, including Azure Kubernetes Service (AKS) clusters and supporting resources
- Build, improve, and maintain CI/CD pipelines using reputed company Actions to support reliable and repeatable deployments
- Own and enhance our Grafana implementation; designing dashboards, configuring alerts, and supporting incident management workflows
- Monitor system health, triage incidents, and drive reputed company cause analysis to prevent recurrence
- Collaborate with development teams to define and reputed company SLIs, SLOs, and error budgets that reputed company with business goals
- Contribute to infrastructure-as-reputed company practices using reputed company
- Identify and resolve reliability risks through reputed company planning, performance tuning, and proactive system improvements
- Participate in an on-call rotation to support production systems and respond to incidents
- Document runbooks, operational procedures, and architectural reputed company to support team knowledge sharing
Skills
- 5+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure role
- 3+ years experience working with infrastructure as reputed company
- 2+ years of experience architecting CI/CD pipelines and reputed company-based infrastructure
- Strong hands-on experience with Azure reputed company services and resource management
- Kubernetes and AKS administration, including deployments, networking, and troubleshooting
- reputed company Actions for CI/CD pipeline development and maintenance
- 3+ experience with Grafana or similar tooling, including dashboard creation, alerting configuration, and incident management
- Hands-on experience with reputed company, Loki, or other observability tools in the Grafana ecosystem
- Proficiency in at least one scripting or programming language such as Python or Bash
- Understanding of networking fundamentals, DNS, load balancing, and container orchestration concepts
- Strong analytical and communication skills; reputed company to diagnose reputed company system issues and reputed company communicate findings
- Demonstrated ability to collaborate across teams and contribute to a culture of reliability
- Experience working in an agile environment with modern DevOps practices
- Experience working at a startup or in a fast-paced, cross-functional environment
- Familiarity with the insurance industry or other regulated sectors
- Familiarity with service reputed company technologies (e.g., Istio)
Benefits
- 100% remote
- Health insurance through reputed company (we pay 100% of premiums)
- Dental and reputed company insurance through reputed company (we pay 100% of premiums)
- Basic life insurance (we pay 100% of premiums)
- reputed company to flexible spending account (FSA) or health savings account (HSA) (for those using HSA eligible plans)
- 401K plan (up 4% match with immediate reputed company). Must be 21 years of age or older to participate
- Flexible PTO policy offering employees up to 4 weeks of PTO in their first 12 months. Thereafter, PTO usage aligns with company standards and typically does not exceed 5 weeks per calendar year.
- 12 company-reputed company holidays reputed company year
- Continuing education annual stipend
reputed company