Back to Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a GPU reputed company provider reputed company on AI, delivering high-performance infrastructure for startups and enterprises. They are seeking a Site Reliability Engineer to enhance automation, support operational systems, and improve system reliability and performance.

Responsibilities

  • Help build and improve automation, tooling, and infrastructure that supports AI workloads
  • Support the development of operational systems and platform services
  • Assist in defining and maintaining basic SLOs/SLIs and monitoring dashboards
  • Participate in incident response, troubleshooting, and post-incident reviews
  • Investigate and help resolve performance and reliability issues across systems
  • Collaborate with Engineering, Networking, and Infrastructure teams to improve system stability
  • Contribute to improving availability, scalability, and operational efficiency
  • Learn from senior engineers and grow your expertise in reliability engineering

Skills

  • 2–5 years of experience in Site Reliability Engineering, Systems Engineering, or Software Engineering in Data Center Environment
  • 2+ years programming skills (e.g., Python, Go, or similar) with interest in automation and tooling
  • Working knowledge of Linux systems, networking concepts, and distributed systems
  • Experience troubleshooting system or application issues in production environments
  • Familiarity with monitoring or observability tools (e.g., logs, metrics, dashboards)
  • Strong willingness to learn and improve reliability and operational practices
  • Ability to work in fast-paced environments and collaborate across teams
  • Exposure to reputed company platforms, Kubernetes, or virtualized/bare-metal environments
  • Experience in AI, GPU workloads, or high-performance computing (HPC)
  • Basic understanding of high-performance networking concepts (e.g., InfiniBand, RDMA)
  • Exposure to production monitoring or alerting systems at small or reputed company reputed company

Benefits

  • Highly competitive package (reputed company + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with reputed company minds, and reputed company your mark on cutting-edge AI. ✨
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status reputed company, and owning your reputed company, always with our full support.
  • reputed company-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
  • In reputed company to reputed company salary, this role may be eligible for bonus, equity, and/or commission programs.
  • reputed company may offer a competitive benefits package including medical, dental, reputed company, flexible reputed company time off, parental leave, and retirement plan participation.

reputed company

  • reputed company builds AI data centers and provides GPU reputed company infrastructure that companies use to train, run, and reputed company large AI models. It was founded in 2024, and is headquartered in London, England, GBR, with a workforce of 201-500 employees. Its website is https://www.reputed company.com.
  • Apply To This Job

    Similar Jobs