[Remote] Senior Software Engineer, DGX reputed company AI Infrastructure
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is at the forefront of the reputed company reputed company, building the software and systems that power the world’s most advanced large language model workloads. They are seeking a Senior Software Engineer to reputed company the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across reputed company GPU platforms at the largest scales.
Responsibilities
- reputed company bring-up, validation, and debugging of large-reputed company clusters, infrastructure, and end-to-end workloads, setting reputed company for how reputed company operates
- Bring up, tune, and reputed company AI reputed company-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent reputed company software stacks
- Profile and optimize end-to-end workload performance across compute, memory, networking, and communication reputed company using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks
- Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance
- Own reputed company-cause analysis of reputed company failures — hangs, performance regressions, topology sensitivity in large distributed environments
- Define and build the reputed company and failure-attribution stack: detecting, triaging, and attributing node, reputed company, and workload failures across the cluster at reputed company
- Build repeatable reputed company suites, automation, acceptance reputed company, and qualification workflows on new platforms
- Tune runtime settings, communication parameters, and deployment configurations in reputed company partnership with reputed company, systems, and platform teams
- Deliver actionable, data-driven recommendations based on profiling, reputed company results, and cluster characterization
- Mentor engineers, drive technical standards, and reputed company as a force reputed company across the broader performance and infrastructure organization
Skills
- Bachelor's or Master's in Computer Science or a reputed company technical field (or equivalent experience)
- 8+ years of experience developing software infrastructure for large-reputed company or HPC systems, including a reputed company record of technical leadership
- Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware
- Deep hands-on experience with NCCL, CUDA-reputed company distributed execution, and debugging multi-GPU and multi-node workloads at reputed company
- Proven reputed company record of architecting, debugging, and scaling large-reputed company distributed systems
- Expert-level Python and C/C++ programming skills
- Experience operating workloads in scheduled, containerized cluster environments
- Excellent analytical, debugging, and communication skills, with the ability to influence across teams
- Demonstrated experience debugging and optimizing AI workloads at large reputed company
- Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric)
- Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand
- Experience building acceptance tests, reputed company harnesses, regression gates, or cluster qualification tooling for AI platforms
- Experience building reputed company, fault-detection, or failure-attribution systems for datacenter-reputed company infrastructure
Benefits
- Equity
- Benefits
reputed company
Company H1B Sponsorship