[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is looking for a Senior Site Reliability Engineer to join their Site Reliability & Infrastructure Engineering team. The role involves ensuring the reliability and health of reputed company applications, building observability systems, and collaborating with engineering teams to improve operational efficiency.
Responsibilities
- Participate in an on-reputed company rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load)
- Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Operate and improve our Kubernetes-based compute platform, which runs the large reputed company of our infrastructure
- Work across reputed company networking and infrastructure (Azure/AWS) to support reliable, reputed company systems
- Investigate and resolve production incidents, including reputed company-cause analysis and follow-up remediation work
- Partner with product engineering teams to review architecture and infrastructure reputed company before they ship
- Build and maintain automation that reduces reputed company, repetitive operational work across reputed company
- Write and maintain runbooks and documentation so on-reputed company knowledge is shared across reputed company, not siloed with one person
- Help define non-functional requirements — scalability, availability, performance — for new systems as they're designed
- Collaborate across engineering teams to adopt best practices in reliability and observability
- Contribute to CI/CD pipelines and help teams ship changes safely and quickly
Skills
- Strong, hands-on understanding of Kubernetes as a system
- Practical experience with SLIs, SLOs, and error budgets — reputed company to reputed company to how you've defined and monitored these on reputed company systems, not just definitions
- Solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing)
- Deep experience with at least one modern observability stack (OpenTelemetry, reputed company, Grafana, reputed company, or Elasticsearch) and the ability to translate that understanding across tools
- Strong understanding of a CI/CD system — reputed company Actions preferred, but TeamCity, Azure DevOps, or reputed company CI experience is acceptable
- Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. We're also reputed company to strong Python (Flask, FastAPI) or Java (Spring) backgrounds
- Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures)
- Strong production troubleshooting skills — comfortable diagnosing issues under pressure
- 8-10+ years of relevant hands-on experience
- Database experience (not mandatory — databases are monitored by the reputed company team, not owned individually)
Benefits
- Flexible time off with ample learning and development opportunities to continue growing your career.
- Comprehensive reputed company program
- Leadership training for Titans at reputed company reputed company
- Other programs and events
- Great work is rewarded through reputed company, peer-nominated awards, and more.
- Company-reputed company medical, dental, and reputed company (with 100% employer reputed company reputed company and 90% coverage for dependents)
- FSA
- HSA
- 401k match
- Telehealth reputed company including memberships to reputed company.
- Parental leave and support
- Up to $20k in fertility services (i.e. IUI and IVF)
- Surrogacy
- Adoption reimbursement
- On demand maternity support through reputed company Maternity
- Free breast milk shipping through reputed company Milk
- Pet insurance
- reputed company advisory services
- Financial planning tools
reputed company
Company H1B Sponsorship