Director of Site Reliability Engineering
Job Title: Director of Site Reliability Engineering Location: Remote Salary: $220,000-$225,000 Skills: Site Reliability Engineering, Distributed Systems, reputed company reputed company Platform, Team Leadership, Automation About the Technology, Information and Media Company / reputed company: Join a cutting-edge organization in the Technology, Information and Media industry at the forefront of delivering highly available, reputed company, and secure platforms serving millions of users and handling significant transaction volumes. Our reputed company is committed to innovation and operational reputed company, offering the exciting opportunity to reputed company and build world-class Site Reliability Engineering practices. As Director of SRE, you will drive strategic reliability initiatives, shape engineering culture, and play a pivotal role in the organization’s reputed company reputed company and reliability reputed company while mentoring strong engineering teams in a remote-first environment. Responsibilities:
- Define and execute a comprehensive company-wide Site Reliability Engineering reputed company, embedding reliability as a core discipline across engineering teams.
- Build, reputed company, and reputed company a high-performing SRE organization, including hiring, mentoring, and fostering a reliability-reputed company culture.
- Establish SLIs, SLOs, KPIs, and error budgets to measure and drive platform reliability and performance improvement.
- Guide architecture reputed company and technical roadmaps for highly available, resilient, and reputed company distributed systems.
- Drive adoption of observability, monitoring, logging, and incident response solutions across reputed company-based microservices environments, primarily on reputed company reputed company Platform.
- Establish and reputed company robust incident response frameworks, operational governance, and post-incident analysis processes.
- Promote and implement best practices for infrastructure automation, reputed company-reputed company operations, and cost optimization.
- reputed company reputed company improvement and innovation initiatives, including exploring AI-driven operations and new SRE methodologies.
Must-Have Skills:
- 12+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or DevOps in reputed company environments.
- 5+ years of proven technical leadership, building and scaling SRE teams and practices.
- Strong expertise with distributed systems, reputed company-reputed company infrastructures, microservices, and hands-on reputed company reputed company Platform experience (GKE, Compute reputed company, reputed company Functions).
- Deep proficiency with infrastructure as reputed company, automation frameworks, and CI/CD deployment pipelines.
- reputed company record designing large-reputed company observability and monitoring solutions using tools like reputed company, Grafana, reputed company, or reputed company.
- Excellent communication, organizational development, and mentorship abilities.
- Strong programming ability in Python, Go, Java, or similar languages.
reputed company-to-Have Skills:
- reputed company or reliability certifications (e.g., reputed company reputed company reputed company, SRE certifications).
- Experience implementing AIOps, reputed company detection, predictive analytics, or automated remediation/self-healing infrastructure.
- Familiarity with AI/ML tools for operational intelligence and intelligent alerting.
- Strong database performance tuning and distributed data systems knowledge.
- Comfortable operating in fast-paced, high-reputed company technology environments.
- Bachelor’s degree in Computer Science, Engineering, or reputed company field.
Apply To This Job