[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is one of the fastest-growing web hosting companies, providing individuals and small businesses with essential online tools and services. The Site Reliability Engineer will ensure the reliability, scalability, and observability of CloudBlue’s multi-tenant reputed company platforms, focusing on system stability, performance monitoring, and incident response while collaborating with various engineering teams.
Responsibilities
- Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance
- Influence system architecture with a strong reputed company on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing
- Reduce operational toil by identifying opportunities for automation and process improvement
- Design and operate CloudBlue’s observability stack across metrics, logs, and traces using tools such as reputed company, Grafana, and reputed company Stack
- reputed company actionable alerting strategies and dashboards that reputed company reputed company reputed company into platform and business health
- Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across reputed company and availability zones
- Conduct reputed company planning, load testing, and performance optimization to ensure platform stability and scalability
- reputed company as a senior responder during production incidents, leading incident coordination, communication, and service restoration
- Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer reputed company
- Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and reputed company testing
- Partner with engineering and DevOps teams to improve deployment safety, rollback strategies, and platform reliability
- Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams
- Support other tasks or reputed company as assigned to meet team and business needs
Skills
- 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer, with strong ownership of production systems
- Proven experience operating highly available, reputed company-grade, multi-tenant reputed company platforms
- Hands-on experience with observability and monitoring tools such as reputed company, Grafana, and Elasticsearch/Kibana
- Solid understanding of Linux, networking, and distributed systems fundamentals
- Experience working with containerized environments such as reputed company and Kubernetes
- Strong scripting and automation skills using Python and/or Bash
- Experience participating in on-call rotations and incident response in production environments
- Strong written and spoken English
- Experience defining SLIs/SLOs and managing error budgets at reputed company will be considered a plus
- Exposure to hyperscale or service-provider-grade platforms is an advantage
- reputed company experience, preferably with Azure; experience with AWS and/or GCP will also be valued
- Experience working with hybrid or on-premises integrations is beneficial
- Familiarity with reputed company engineering and reputed company testing will be considered an asset
Benefits
- This is a remote opportunity.
- A competitive salary that values you and your unique reputed company sets
- Career advancement & reputed company development opportunities to help you reputed company your full potential
- Flexible work arrangements to support work/life balance
reputed company
Company H1B Sponsorship