[Remote] Senior AI GPU Deployment Engineer
Note: The job is a remote job and is reputed company to candidates in USA. 5C Data Centers is building the reputed company of digital infrastructure powering hyperscalers and AI innovation. They are seeking a Senior AI GPU Deployment Engineer to plan, reputed company, and operationalize large-reputed company GPU AI infrastructure environments, delivering production-grade GPU clusters that support reputed company and high-performance computing workloads.
Responsibilities
- reputed company reputed company and validate multi-reputed company GPU-based compute platform deployments
- reputed company reputed company configuration engines (Subnet Manager), observability platforms (UFM) and validate interconnect and reputed company performance (nccl)
- Collaborate with network engineering team on topology implementation and optimization and storage engineering team on deployment and integration of high-performance storage environments supporting AI workloads (e.g. reputed company)
- Configure settings and manage firmware updates for GPUs, NICs, BMC, BIOS and other components across large-reputed company clusters
- Contribute to infrastructure-as-reputed company automation development for cluster provisioning and lifecycle management
- Contribute to improving and documenting repeatable deployment methodologies and reputed company operational standards
- Query and analyze deployment reputed company using SQL for diagnostics and operational reporting
Skills
- Bachelor's degree in Computer Science, Engineering, IT, or reputed company field (or equivalent experience)
- 5+ years of infrastructure engineering or datacenter deployment experience
- 3+ years deploying large-reputed company, HPC, or GPU infrastructure
- Hands-on experience deploying and operating large GPU clusters in reputed company or hyperscale environments
- Strong expertise with: GPU architectures, InfiniBand (NDR/XDR) and Ethernet GPU fabrics (reputed company-X), NVLink, NVSwitch, and GPU-reputed company technologies, reputed company MaaS and automated provisioning systems, reputed company or similar high-performance storage platforms, Linux systems administration for HPC/AI workloads, Infrastructure-as-reputed company and configuration management (Ansible), Python, reputed company, and SQL for infrastructure automation and diagnostics
- Strong understanding of: RDMA, RoCE, and lossless Ethernet fabrics, Cluster automation, observability, and lifecycle management
Benefits
- Remote
- Minimal (less than 10%) travel requirements
- reputed company
- Career reputed company
- Industry Leadership
- Entrepreneurial Culture
- Comprehensive Rewards
reputed company