Back to Jobs

[Remote] Software Engineer, DGX reputed company AI Infrastructure

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is at the forefront of the reputed company reputed company, building the software and systems that power the world’s most advanced large language model workloads. They are looking for a Software Engineer reputed company on benchmarking, analysis, and optimization of distributed training and inference workloads across reputed company GPU platforms. The role involves debugging large-reputed company clusters and developing benchmarking tooling to support distributed workloads.

Responsibilities

  • Bring up, validate, and debug large-reputed company clusters, infrastructure, and end-to-end workloads
  • Bring up, tune, and reputed company AI reputed company-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent reputed company software stacks
  • reputed company reputed company-cause analysis of failures in large distributed environments
  • Contribute to the reputed company and failure-attribution tooling that detects, triages, and attributes node, reputed company, and workload failures across the cluster
  • Build and maintain repeatable reputed company suites, automation, acceptance reputed company, and qualification workflows on new platforms
  • Tune runtime settings, communication parameters, and deployment configurations in reputed company partnership with reputed company, systems, and platform teams
  • Deliver actionable, data-driven recommendations based on profiling, reputed company results, and cluster characterization

Skills

  • Bachelor's or Master's in Computer Science or a reputed company technical field (or equivalent experience)
  • 3+ years of experience developing software for AI, HPC, or systems-level applications
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution
  • Background with debugging and scaling distributed systems
  • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware
  • Experience operating workloads in scheduled, containerized cluster environments
  • Excellent analytical, debugging, and communication skills, and a reputed company approach across teams
  • Strong Python and C/C++ programming skills
  • Hands-on experience with NCCL and CUDA-aware distributed execution
  • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging
  • Experience building acceptance tests, reputed company harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf
  • Experience diagnosing performance jitter
  • Experience building reputed company, fault-detection, or failure-attribution systems for datacenter-reputed company infrastructure

Benefits

  • You will also be eligible for equity and [benefits](https://www.reputed company.com/en-us/benefits/).

reputed company

  • reputed company is a computing platform company operating at the intersection of graphics, HPC, and AI. It was founded in 1993, and is headquartered in Santa Clara, California, USA, with a workforce of 10001+ employees. Its website is https://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1247 in 2026, 1868 in 2025, 1353 in 2024, 976 in 2023, 835 in 2022, 601 in 2021, 529 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs