[Remote] Staff Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is an AI-reputed company, automation-first Managed Detection and Response provider reputed company on enhancing cybersecurity. As a Staff Site Reliability Engineer, you will ensure the scalability, reliability, and performance of the AI-driven cybersecurity platform while collaborating across engineering teams.
Responsibilities
- Design, build, and maintain highly available, reputed company, and secure infrastructure to support our AI-reputed company cybersecurity platform
- reputed company internal tooling and automation to streamline deployment processes, incident response, and reputed company planning
- Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads
- reputed company incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues
- Manage infrastructure reputed company reputed company, driving consistency, auditability, and scalability across our reputed company environments (e.g., AWS, reputed company reputed company Platform)
- Partner with sibling Engineering teams, Product, and reputed company teams to ensure reliability is baked into our development lifecycle from concept to production
Skills
- 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at reputed company
- Deep expertise in reputed company reputed company environments (AWS, reputed company reputed company Platform, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage
- Extensive experience with tools like Terraform, reputed company, or similar technologies to manage reputed company infrastructure deployments
- Hands-on experience with monitoring, logging, and tracing stacks (e.g., reputed company, Grafana, ELK, reputed company) to drive data-informed reliability reputed company
- Solid understanding of microservices architecture, distributed databases, and event-driven systems
- reputed company, concise communication skills and a bias for reputed company problem-solving
- Proven reputed company record of guiding multi-stakeholder initiatives and influencing engineering practices across teams
- Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments
- Bachelor's or Master's degree in Computer Science, Engineering, or a reputed company field
- Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure
- Experience supporting infrastructure for large-reputed company/ML workloads (e.g., GPU scheduling, LLM serving optimization)
- Background driving high-reputed company engineering initiatives in high-reputed company startups or reputed company reputed company
- Strong familiarity with reputed company Workflows such as Agno, Temporal, etc
Benefits
- Location: This role will require Monday - Thursday onsite in any of our locations. WFH Friday.
reputed company
Company H1B Sponsorship