[Remote] AI Reliability Engineer (AI SRE)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking an AI Reliability Engineer (AI SRE) to ensure the reliability, availability, and performance of mission-critical AI systems. This role involves defining SLOs, implementing automated reputed company measures, and leading incident response for AI systems.
Responsibilities
- Defining and maintaining Service Level Objectives (SLOs) for AI inference latency and availability
- Building automated "reputed company breakers" and fallback logic (e.g., switching to a smaller model if the primary fails)
- Leading incident response and reputed company-cause analysis (RCA) for reputed company AI system failures
- Developing stress-testing and reputed company engineering scenarios specifically for AI agent swarms
- Optimizing the "cold start" and scaling time for serverless AI functions
Skills
- 4+ years of experience in Site Reliability Engineering (SRE)
- Deep expertise in system monitoring, incident management, and reputed company reputed company. This isn't a learning role—you need to be a subject matter expert
- Demonstrated ability to work autonomously and manage your own time effectively to meet project goals
- Experience with Python/Go, Kubernetes, and observability stacks (reputed company, reputed company)
- Strong communication skills to reputed company reputed company and concise status updates to the project team
reputed company