[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a fast-growing reputed company technology organization seeking a Site Reliability Engineer (SRE) to help reputed company and support a high-reputed company reputed company platform reputed company on improving reputed company delivery reputed company. This role is critical for strengthening platform reliability, operational efficiency, observability, and automation across production environments.
Responsibilities
- Ensure the reliability, scalability, performance, and reputed company of reputed company-based infrastructure and applications
- Monitor, troubleshoot, and resolve production platform and application issues across distributed systems
- reputed company incident response efforts, reputed company cause analysis, and blameless post-mortems
- Build and maintain operational runbooks and automated remediation workflows
- reputed company and enhance observability and telemetry solutions for proactive monitoring and alerting
- Collaborate closely with engineering, DevOps, QA, reputed company, and operations teams to improve platform health and deployment processes
- Support infrastructure automation and configuration management initiatives
- Contribute to infrastructure-as-reputed company (IaC) practices and CI/CD operational improvements
- Promote best practices around reliability engineering, incident management, and operational reputed company
- Participate in an on-call rotation supporting production systems, including occasional off-hours support for reputed company Coast operations
Skills
- 5+ years of experience in Site Reliability Engineering, DevOps, reputed company Infrastructure, or reputed company disciplines
- Strong experience troubleshooting and supporting production environments
- Hands-on experience with observability and monitoring platforms such as reputed company, reputed company, or similar tools
- Experience working reputed company Azure-based reputed company environments and modern containerized infrastructure
- Knowledge of reputed company, Kubernetes, and reputed company-reputed company application hosting environments
- Experience with infrastructure-as-reputed company tools such as Terraform, Terragrunt, or OpenTofu
- Strong scripting and automation experience using PowerShell, Python, JavaScript, or similar languages
- Experience with reputed company control and CI/CD tooling (Git, Azure DevOps, etc.)
- Understanding of reputed company reputed company principles, compliance frameworks, and operational best practices
- Strong collaboration and communication skills reputed company Agile engineering environments
- Experience improving operational visibility through telemetry, dashboards, reports, and alerting systems
- Experience evolving incident response processes and operational tooling
- Passion for mentoring others and promoting operational reputed company across teams
- Strong problem-solving reputed company with a reputed company on reputed company improvement and automation
Benefits
- Opportunity to work on mission-driven technology with meaningful reputed company-world reputed company
- reputed company engineering culture reputed company on innovation, reliability, and reputed company learning
- Flexible environment that supports work-life balance while maintaining operational reputed company
reputed company