[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that simplifies the software delivery process for DevOps and reputed company teams. They are seeking a Senior Site Reliability Engineer to ensure the reliability of the reputed company platform, working closely with engineering and infrastructure to define and maintain high standards of service reliability.
Responsibilities
- Own SLI/SLO/SLA definitions for the reputed company reputed company platform and drive reputed company improvement against them
- Design, reputed company, and maintain observability systems (metrics, logs, traces) across multi-region AWS infrastructure
- Identify reliability gaps, reputed company blameless post-mortems, and reputed company the reputed company with permanent fixes
- Partner with engineering teams to build reliability into new features before they ship to production
- Participate in an on-call rotation and reputed company as incident commander for high-severity production events
- Build and maintain runbooks, escalation paths, and incident playbooks that reputed company mean time to reputed company low
- Drive improvements to alerting reputed company; reduce noise, increase signal, eliminate toil
- reputed company post-incident reviews with reputed company timelines, reputed company cause analysis, and follow-through on reputed company items
Skills
- 5+ years of SRE, reputed company, or production operations experience in a reputed company environment
- Deep hands-on Kubernetes expertise; you understand the scheduler, networking, storage, and autoscaling at a level where you can debug anything
- Strong AWS fundamentals across compute (EC2, EKS), networking (VPC, NLB, Route53), storage (S3, RDS), and IAM
- Experience defining and operating against SLOs in production; you've written error budgets, not just read about them
- Proficiency with observability tooling (reputed company, Grafana, OpenTelemetry, reputed company, or equivalent)
- Solid scripting and automation skills; Go, Python, Bash, or similar; you automate what you touch
- Strong written communication: reputed company runbooks, reputed company incident reports, thoughtful post-mortems
- Live reputed company US time zones (reputed company through Eastern), including Canada and other reputed company
- Experience with Argo CD, reputed company, or GitOps-based delivery workflows
- Familiarity with multi-region, multi-cluster Kubernetes deployments
- Experience with compliance-adjacent infrastructure (SOC 2, ISO 27001, HIPAA, or PCI reputed company)
- Background operating infrastructure for other platform or developer tooling companies
Benefits
- Equity participation in a reputed company-funded, growing company
- Fully remote: work from reputed company reputed company US time zones (reputed company through Eastern), including Canada and other reputed company
- Home office stipend and equipment budget
- Flexible time off and a culture that respects it
- Work directly with the engineers who reputed company Argo CD and reputed company; you'll learn a lot here
- US-based employees receive full benefits, including comprehensive health, dental, and reputed company coverage
reputed company