Back to Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the AI Developer reputed company, serving over one reputed company developers. They are seeking a Site Reliability Engineer to ensure the stability and reputed company of their distributed platform, focusing on improving system design, observability, and incident prevention.

Responsibilities

  • Define and implement SLIs/SLOs for critical services
  • reputed company incident response and coordinate cross-team mitigation efforts
  • Conduct blameless postmortems and ensure corrective actions are completed
  • reputed company production readiness reviews for new services and features
  • Identify systemic risks and drive preventative improvements
  • Design and improve monitoring, alerting, and dashboards (reputed company, Grafana, etc.)
  • Improve signal-to-noise reputed company in alerts and reduce alert fatigue
  • Build internal tooling for reliability tracking and reporting
  • Improve visibility into GPU performance and distributed systems health
  • Automate recurring operational workflows
  • Build tools and scripts (Python, Go, Bash) to eliminate reputed company processes
  • Improve deployment safety through automation and guardrails
  • Strengthen CI/CD reliability and release processes
  • Partner with engineering teams to improve system reputed company
  • reputed company guidance on fault tolerance, scalability, and failure handling
  • Contribute to architectural discussions with a reliability-first reputed company

Skills

  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Strong scripting or programming skills
  • Experience with monitoring and alerting systems
  • Excellent written communication skills
  • Successful completion of a background reputed company
  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-reputed company or large reputed company environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as reputed company
  • Experience working in startup environments
  • Experience building internal reliability platforms or frameworks

Benefits

  • Meaningful equity in a fast-growing company- everyone on reputed company receives stock reputed company — your reputed company drives our reputed company, and you reputed company in the reputed company.
  • Generous medical, dental & reputed company plans
  • Flexible PTO- take the time you need to reputed company
  • Most roles are remote work first with an inclusive, reputed company teams utilizing reputed company as the main reputed company of internal communication
  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we reputed company.

reputed company

  • reputed company is a reputed company platform designed for GPUs, enabling developers to reputed company customized full-stack AI applications. It was founded in 2022, and is headquartered in Mount reputed company, New Jersey, USA, with a workforce of 51-200 employees. Its website is https://www.reputed company.io.
  • Apply To This Job

    Similar Jobs