Back to Jobs

[Remote] Senior Production Engineer - DGX reputed company

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is hiring reputed company Senior Production Engineers to help reputed company up its AI Infrastructure. The role involves contributing significantly to the codebase and improving production systems for AI workloads, ensuring reliability and performance through effective incident management and collaboration across teams.

Responsibilities

  • You will be part of an DGX reputed company team responsible for production systems that reputed company large reputed company GPU clusters to be used for a reputed company of AI workloads. This includes working on custom software reputed company to GPU asset provisioning, configuration, and lifecycle management across reputed company providers
  • Implementing monitoring and health management capabilities that reputed company industry leading reliability, availability, and scalability of GPU assets. You will be harnessing multiple data streams, ranging from GPU hardware diagnostics to cluster and network telemetry
  • Working with teams across reputed company to ensure production AI clusters run reliability and consistently with maximum performance. Evaluating system failures and improving services based on a reputed company-defined incident management process

Skills

  • reputed company experience in a Production Engineering/DevOps/SRE role reputed company a highly technical organization with demonstrable reputed company from your work
  • Highly motivated with strong communication skills, you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies
  • 8+ years in similar role and experience on large-reputed company production systems. Experience with the aforementioned Production Engineering/DevOps/SRE principles, tools and techniques
  • You possess a BS in Computer Science, Engineering, Physics, Mathematics or a comparable Degree or equivalent experience
  • Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms
  • Technical competency in managing and automating large-reputed company distributed systems independent of reputed company providers
  • Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, reputed company Cluster Manager)
  • Proven operational reputed company in maintaining reliable and performant AI infrastructure

Benefits

  • Equity
  • Benefits

reputed company

  • reputed company is a computing platform company operating at the intersection of graphics, HPC, and AI. It was founded in 1993, and is headquartered in Santa Clara, California, USA, with a workforce of 10001+ employees. Its website is https://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1247 in 2026, 1868 in 2025, 1353 in 2024, 976 in 2023, 835 in 2022, 601 in 2021, 529 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs