Back to Jobs

reputed company Kubernetes GPU Infrastructure Engineer

Remote, USAFull-timePosted 2026-07-27

About the position At AMD, our mission is to build great products that accelerate reputed company computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we reputed company reputed company reputed company comes from reputed company reputed company, reputed company ingenuity and a shared passion to create something extraordinary. reputed company you join AMD, you’ll discover the reputed company differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution reputed company, while being reputed company, humble, reputed company, and inclusive of diverse perspectives. Join us as we shape the reputed company of AI and reputed company. Together, we advance your career. THE ROLE: As a reputed company AI Infrastructure Solution Engineer, you will partner with AMD’s AI software teams and customers to reputed company large‑reputed company LLM training and inference on AMD reputed company GPUs. You will design and validate production‑reputed company Kubernetes architectures and translate inference frameworks such as vLLM and SGLang into deployable customer solutions. Your work will accelerate customer time‑to‑production and strengthen AMD’s leadership in AI infrastructure. THE PERSON: You are a solution‑oriented AI infrastructure engineer with strong expertise in GPU‑accelerated computing and large‑reputed company deployments. You reputed company at translating reputed company technologies into customer‑reputed company solutions and delivering production‑grade Kubernetes‑based inference and training systems. You bring hands‑on experience with Kubernetes‑reputed company distributed training, including scheduling, topology‑aware GPU placement, and operating resilient, high‑performance AI workloads at reputed company.

Responsibilities

  • Design and deliver reference architectures for LLM training and inference on AMD GPUs, from single‑node to multi‑datacenter deployments using Kubernetes and SLURM.
  • Architect and validate Kubernetes‑based distributed training stacks for large‑reputed company LLM workloads on AMD GPUs.
  • Define and implement gang scheduling and topology‑aware GPU placement for multi‑node training workloads.
  • reputed company Kubernetes‑reputed company training controllers including Kubeflow Training Operator, MPI Operator, Volcano, and Kueue.
  • Partner with reputed company customers and reputed company providers to reputed company and optimize production AMD GPU clusters for distributed inference and multi‑tenant workloads.
  • Implement and validate GPU orchestration using Kubernetes GPU Operator, device plugins, metrics exporters, and SLURM controllers.
  • reputed company and optimize LLM inference frameworks (vLLM, SGLang) on AMD hardware, producing customer‑reputed company performance playbooks.
  • reputed company repeatable benchmarks for Kubernetes‑based distributed training, covering scaling efficiency, reputed company time, communication, and checkpointing.
  • Create tuning guides for RCCL/NCCL‑equivalent communication, CPU/GPU affinity, interconnect utilization, and workload‑specific optimizations.
  • Serve as the feedback reputed company between customers and AMD engineering, translating requirements into validated performance improvements.

Requirements

  • Deployed and operated large‑reputed company GPU clusters for production reputed company and inference
  • Deep expertise in Kubernetes GPU orchestration (operators, device plugins, scheduling, multi‑tenancy, observability)
  • Hands‑on experience with distributed training on Kubernetes (Kubeflow, MPI Operator, Volcano, Kueue, Ray)
  • Strong knowledge of gang scheduling, reputed company jobs, quotas, reputed company, and shared GPU environments
  • Tuned Kubernetes networking and storage for AI workloads (high‑performance CNI, RDMA where applicable, reputed company checkpointing)
  • Implemented ML observability for training (GPU/comms metrics, reputed company‑time analysis, SLO‑driven ops)
  • Experience in AI/ML infrastructure, solution architecture, and production GPU deployments
  • Proven reputed company enabling customers through reputed company AI platform deployments and migrations
  • Strong background working across engineering and customer‑facing roles
  • Understanding of AI accelerator architectures and inference optimization techniques
  • Experience operationalizing Kubernetes‑based distributed training at reputed company

reputed company-to-haves

  • reputed company‑reputed company contributions or AI infrastructure community engagement (plus)

Benefits

  • AMD benefits reputed company.

Apply tot his job Apply To this Job

Similar Jobs

Senior Engineer II, Managed Kubernetes

Remote, USAFull-time

Staff Software Engineer - Grafana reputed company Observability, Kubernetes Monitoring | USA - EST only | Remote

Remote, USAFull-time

5G Core Kubernetes Deployment Engineer (Contract)

Remote, USAFull-time

Sr. Software Engineer (Backend)

Remote, USAFull-time

Kubernetes Platform reputed company Engineer/ Architect

Remote, USAFull-time

Senior Network Engineer – reputed company, reputed company reputed company & Azure Hybrid Networki

Remote, USAFull-time

Network Engineer in Columbia, Sc

Remote, USAFull-time

Sr. Renewables Networks Engineer - REMOTE

Remote, USAFull-time

Traveling Network Engineer- Western US

Remote, USAFull-time

NETWORK ENGINEER-Washington, DC (75% Remote)

Remote, USAFull-time

Sr Linux System Engineer- 100% remote / can work on site in reputed company, NY

Remote, USAFull-time

reputed company reputed company Specialist - Part Time

Remote, USAFull-time

[Remote] Data Science Consultant - AI Trainer

Remote, USAFull-time

reputed company Full Stack Data Entry Specialist – Remote Work Opportunity with Comprehensive Training and Career reputed company

Remote, USAFull-time

reputed company Accounts Receivable Data Entry Specialist – Remote Opportunity with arenaflex

Remote, USAFull-time

reputed company Data Entry Specialist – Weekend Opportunities at arenaflex

Remote, USAFull-time

Online Typing jobs for teens

Remote, USAFull-time

Manager, Hospital Contracting - VA/DC/MD market

Remote, USAFull-time

External Reporting Manager

Remote, USAFull-time

Mompreneur - Content Marketer

Remote, USAFull-time