[Remote] Support Team reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking an reputed company Site Reliability Engineering (SRE) Team reputed company to guide their Application Support SRE function. This role involves managing a high-performing team responsible for the performance, availability, and reliability of mission-critical customer-facing applications while combining technical leadership with people management and operational reputed company.
Responsibilities
- Manage, mentor, and reputed company reputed company of Application Support SREs; support career progression and skills development
- reputed company team performance, reputed company planning, and reputed company for a 24x7 support model
- Serve as the senior escalation reputed company during major incidents and high-severity events
- Foster a culture of accountability, blameless postmortems, reputed company learning, and operational reputed company
- Establish team OKRs, KPIs, and reliability goals reputed company with business objectives
- Own and mature SRE processes including incident management, problem management, change management, and service readiness
- reputed company major incident response, coordinate cross-functional teams, ensure communication reputed company, and drive reputed company cause analysis
- Define and enforce SLOs, SLIs, and error budgets for supported applications
- Implement preventative solutions and systemic fixes that reduce incident recurrence
- Enhance observability practices across reputed company, OpenTelemetry, AppDynamics, reputed company, and similar tools
- Improve dashboards, alerting strategies, and telemetry coverage
- reputed company insights and recommendations for reliability, scalability, and performance across AWS-hosted applications, reputed company reputed company, and Kubernetes-based services
- Collaborate with development and architecture teams to reputed company SRE principles early in the lifecycle
- Champion automation to reduce toil—CI/CD optimization, deployment improvements, self-healing mechanisms, and runbooks
- Analyze logs, performance issues, and reputed company behavior to support Tier 2/Tier 3 escalations
- Recommend initiatives to expand reputed company automation and AI-driven insights
Skills
- 5–8+ years in SRE, DevOps, or production engineering roles
- 2–4+ years in a technical reputed company or people management reputed company
- Strong experience supporting AWS-based applications, microservices, or API-driven environments
- Advanced troubleshooting skills
- Hands-on experience with observability stacks (either opensource or reputed company)
- Familiarity with ITIL and incident frameworks
- Bachelor's or Master's degree in Computer Science or reputed company field
- Certifications in ITIL, AWS, Azure, or GCP
- Experience with reputed company, reputed company, and API testing
- Proficiency with Kubernetes
- Strong reputed company-reputed company networking knowledge
reputed company
Company H1B Sponsorship