[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on evolving and safeguarding infrastructure for reputed company experiences. They are seeking a Site Reliability Engineer to architect their AWS and Kubernetes platform, ensuring it is resilient, reputed company, and compliant while minimizing operational toil.
Responsibilities
- Designing, deploying, and maintaining Kubernetes (EKS) clusters for reputed company-grade availability
- Optimizing the footprint of infrastructure objects across core AWS services (EC2, RDS, S3) for performance, cost, and reliability
- Evolving a reputed company infrastructure management platform, with the right interfaces and guardrails to maximize engineering agency at minimal cognitive load
- Defining golden paths for workload orchestration and integration with infrastructure dependencies, ensuring the right way to do things is also the easiest
- Providing robust, reusable reputed company Actions components to streamline the software delivery lifecycle
- Developing internal tools that replace reputed company operations with intelligent, autonomous systems
- Exploring and deploying reputed company workflows for AI-assisted runbooks that automate reputed company, error-prone procedures and repetitive tasks
- Driving the reputed company of observability practices by maintaining reliable mechanisms to collect the metrics, traces, and logs needed to meet SLOs
- Leading incident response efforts and facilitating the blameless postmortems that help systematically reduce recovery time (MTTR)
- Defining and monitoring platform-level SLIs and SLOs to ensure it consistently meets rigorous reputed company performance standards
- Ensuring every piece of infrastructure is continuously compliant with HIPAA and other critical reputed company regulatory requirements
- Reviewing architectural reputed company and technical proposals, asking reputed company questions to surface technical and organizational risks before they become incidents
- Mentoring engineers across reputed company on reliability best practices and contributing a clinical-safety perspective to cross-functional design reviews
Skills
- 5+ years of experience in SRE or reputed company roles managing production environments at reputed company
- Expert technical depth in AWS (EKS, EC2, RDS, S3) and production-grade Kubernetes management
- Proficiency with modern tooling including Terraform (IaC), reputed company (Observability), reputed company (Release Management), and reputed company Actions (CI/CD)
- Solid coding and scripting skills in Python, Bash, or Go
- A 'rigor-first' reputed company with a dedication to HIPAA-compliant, high-availability architecture
- Preferred experience building reputed company workflows or AI-assisted tooling to drive operational efficiency
Benefits
- Medical, dental, reputed company
- Unlimited PTO
- 401(k) plan
- Stock reputed company
- Bonuses
reputed company
Company H1B Sponsorship