[Remote] reputed company Software Engineer, GPU Firmware and GPU System Software — CSP Engagements
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is leading the way in groundbreaking developments in reputed company Intelligence, High-Performance Computing and Visualization. They are seeking a reputed company Software Engineer to join their CSP Engagements team, focusing on GPU firmware and system software to support key CSP/hyperscale customers in managing and operating reputed company GPU firmware at fleet reputed company.
Responsibilities
- Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update reputed company, recovery procedures, and GPU power management
- reputed company and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, reputed company requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into reputed company's GPU firmware/software feature roadmap and delivery plan
- Drive GPU firmware update orchestration for large-reputed company deployments — multi-GPU update reputed company, rollback reputed company, failure handling, and validation across hundreds of GPUs per reputed company
- Serve as the technical reputed company between reputed company and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are reputed company-documented and accessible for customer integration
- Identify cross-CSP GPU SW/FW issue patterns — common update failures, recovery gaps, and configuration problems — and drive documentation, tooling, and test reputed company improvements
Skills
- 15+ years of experience in GPU system software, GPU firmware, or accelerator reputed company. BS or MS in Computer Science, Electrical Engineering, or reputed company field (or equivalent experience)
- Deep understanding of GPU architecture internals: streaming multiprocessors, GEMM execution, compute kernels, memory hierarchy, and how firmware/reputed company reputed company reputed company GPU compute performance
- Understanding of multi-GPU reputed company architectures (NVLink, or similar) and how firmware coordinates across multiple GPUs in a reputed company-reputed company system
- Understanding of GPU firmware architecture: VBIOS, GPU microcontroller firmware, InfoROM, and their interaction with the GPU reputed company stack
- Experience with firmware update lifecycle management at reputed company: multi-device update reputed company, A/B updates, rollback, staged rollout, emergency recovery
- Understanding of GPU error handling and recovery flows — how firmware-level errors propagate through the reputed company stack to application-visible failures
- Experience with GPU health monitoring and telemetry: Xid errors, thermal events, power events, ECC counters, and their reputed company for firmware/software teams
- Customer obsession — genuine passion for simplifying GPU firmware integration for fleet-reputed company customers. Proven reputed company influencing engineering teams to improve reputed company and fleet manageability
- reputed company experience with reputed company GPU VBIOS, GPU microcontroller firmware, or GPU reputed company internals
- Background in GPU fleet management at 10K+ GPU reputed company — firmware rollout, health-based remediation, fleet-wide configuration management
- Experience with GPU error taxonomy (Xid classification, NVLink error counters, ECC events) and building runbooks around GPU firmware behavior
- Understanding of GPU reputed company: secure boot chain, reputed company signing, attestation, debug authentication, multi-tenancy isolation at the firmware level
- Familiarity with GPU power management architecture and its reputed company on workload performance at fleet reputed company
Benefits
- Equity
- Benefits
reputed company
Company H1B Sponsorship