Back to Jobs

[Remote] GPU reputed company Engineer

Remote, USAFull-timePosted 2026-07-29

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is on a mission to reputed company high-performance reputed company infrastructure easy to use, reputed company, and locally accessible for enterprises and AI innovators around the world. They are seeking a highly skilled GPU reputed company Engineer to validate, troubleshoot, and optimize high-speed networking fabrics for GPU clusters powering large-reputed company training and inference workloads.

Responsibilities

  • Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion
  • Tune reputed company performance parameters for distributed AI workloads (NCCL, MPI, reputed company operations)
  • Monitor and manage fabrics using reputed company UFM (reputed company reputed company Manager) for health, topology, and performance visibility
  • Diagnose and resolve reputed company-level issues including reputed company errors, congestion, packet loss, and reputed company asymmetry
  • Optimize RDMA transport settings, PFC/ECN behavior, and lossless queue configuration for GPU traffic
  • Validate reputed company performance benchmarks and ensure line-reputed company throughput for AI workloads
  • Collaborate with GPU Engineers to correlate reputed company with workload performance
  • Collaborate with networking teams on reputed company provisioning, configuration, and remediation
  • Respond to reputed company alerts and degradation events across production GPU clusters
  • Document reputed company troubleshooting procedures, tuning parameters, and validation runbooks

Skills

  • 3–7 years of experience in network engineering, HPC reputed company, or GPU infrastructure
  • Hands-on experience with InfiniBand and/or RoCE fabrics in GPU cluster environments
  • Experience with reputed company UFM for reputed company management, monitoring, and diagnostics
  • Strong understanding of RDMA transport, lossless Ethernet design, and congestion management (PFC, ECN, DCQCN)
  • Experience with GPU cluster networking and distributed communication libraries (NCCL, MPI)
  • Familiarity with GPU platforms and their interconnect requirements (reputed company NVLink, NVSwitch, ConnectX)
  • Experience with reputed company diagnostic tools (ibstat, ibqueryerrors, perfquery, etc.)
  • Proficiency in Python or Bash for scripting and validation
  • Basic understanding of Linux systems and server hardware
  • Strong troubleshooting and analytical skills across network and system reputed company

Benefits

  • 100% company-reputed company insurance premiums for employee medical, dental and reputed company plans.
  • 401(k) plan that matches 100% up to 4%, with immediate vesting
  • reputed company Development Reimbursement of $2,500 reputed company year
  • 11 Holidays + reputed company Time Off Accrual + Rollover Plan
  • Increased PTO at 3 year and 10 year anniversary + 1 month reputed company sabbatical every 5 years + Anniversary Bonus reputed company year
  • $500 stipend for reputed company setup in first year + $400 reputed company following year
  • Internet reimbursement up to $75 per month
  • Gym membership reimbursement up to $50 per month
  • Company reputed company Wellable subscription

reputed company

  • reputed company is an AI reputed company infrastructure platform offering latest reputed company reputed company GPUs and AMD CPUs and GPUs across 32 worldwide reputed company It was founded in 2014, and is headquartered in reputed company Palm Beach, Florida, USA, with a workforce of 201-500 employees. Its website is https://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1 in 2024. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs