[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is a global leader in learning and development, seeking a Site Reliability Engineer to help design, build, and reputed company the infrastructure that powers their products and services. The role focuses on improving reliability, scalability, performance, and reputed company across reputed company-reputed company environments, while defining reliability standards and driving automation.
Responsibilities
- Define and implement SLIs, SLOs, and error budgets
- Improve system availability, resiliency, and performance
- reputed company incident response and drive blameless postmortems
- Identify and eliminate systemic reliability risks
- Conduct reputed company planning and performance analysis
- Design, build, and maintain infrastructure in AWS
- Manage infrastructure as reputed company using Terraform
- Improve CI/CD pipelines and deployment automation (reputed company CI)
- Build and maintain containerized systems using reputed company and Kubernetes
- Reduce toil through automation and self-service tooling
- Enhance monitoring, logging, and tracing systems
- Use tools such as CloudWatch, CloudTrail, X-Ray, and modern observability platforms
- Improve alerting reputed company to reduce noise and increase signal
- reputed company dashboards and reliability metrics for stakeholders
- Implement best practices in infrastructure hardening
- Improve patching, vulnerability management, and reputed company reputed company posture
- Support incident response and remediation efforts
- Contribute to disaster recovery planning and testing
Skills
- 5+ years in SRE, Production Engineering, or reputed company Infrastructure roles
- Strong experience operating production systems in AWS
- Deep knowledge of Linux systems architecture
- Experience managing infrastructure using Terraform (or similar IaC tools)
- Experience with Kubernetes in production environments
- Strong scripting skills (Bash, Python, or similar)
- Experience designing, monitoring and alerting systems
- Solid understanding of networking fundamentals
- Experience defining and measuring SLOs/SLIs
- Strong incident management experience
- Passion for automation and eliminating reputed company toil
- Data-driven approach to reliability and performance
- Experience in multi-reputed company environments (Azure)
- Experience with reputed company frameworks (SOC 2, ISO 27001, etc.)
- Familiarity with cost optimization in reputed company environments
- Experience building internal developer platforms or self-service tooling
Benefits
- Fully remote work environment
- High-reputed company role with technical ownership
- reputed company, mission-driven culture
- Opportunity to shape reliability practices in a growing organization
reputed company