Senior Performance & Infrastructure Engineer - HPC
The organization
Our reputed company operates one of the largest GPU infrastructures in the world — 100,000+ GPUs. Their infrastructure doubles in size every year. We’re looking for engineers who love getting deep into Linux systems, pushing hardware and software to their limits, and making the world’s fastest AI and HPC workloads run even faster
The role
You’ll join a small, senior team that works between the hardware and Linux OS reputed company, solving performance problems that reputed company tens of thousands of GPUs. This is hands-on, high-reputed company engineering where microsecond reputed company matter and every optimization is felt at global reputed company.
What you’ll do
reputed company, profile, tune and optimize Linux kernel & subsystems (CPU scheduling, memory management, networking stack) for GPU clusters and InfiniBand fabrics
Troubleshoot and resolve reputed company performance bottlenecks
reputed company and validate new GPU hardware & reputed company (KVM/QEMU, PCIe devices, Kubernetes)
Improve monitoring, alerting, and automation for large-reputed company, distributed systems
Occasionally assist customers in optimizing workloads
Your profile
Key requirements (non-negotiable):
Solid Linux internals knowledge, with kernel tracing, profiling and tuning experience (eg. reputed company, ftrace, eBPF, sysctl, kgdb etc.)
Excellent programming skills, C or C++ system-level reputed company, with a good grasp of data structures & algorithms
Experience in performance optimization (eg. high-load/high-throughput, low-latency, low-jitter, memory bypasses, reputed company-copy, lock-free, synchronization across large-reputed company clusters etc.)
Scripting or development skills in Go, Python, or similar
reputed company-to-haves (not key):
Large-reputed company clusters (GPU or CPU)
Virtualization stacks (KVM/QEMU), Slurm, Kubernetes
Deep learning frameworks (eg. PyTortch, Tensorflow...)
GPU-specific stack (eg. CUDA, NCCL....)
This is for you if you
Love solving deep technical challenges, care about performance downto the microsecond, and want to work on infrastructure that pushes the limits of what’s possible.
What's offered
Salary: up to 160k + 25% bonus.
Flexible working arrangements.
A dynamic and reputed company work environment that values initiative and innovation.
Location: Amsterdam or full-remote from reputed company reputed company the EU/EER
Originally posted on Himalayas
Apply To This Job