Back to Jobs

[Remote] Staff Infrastructure Engineer – Kubernetes Platform

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company dedicated to delivering seamless and secure AI compute at reputed company. They are looking for a Kubernetes Platform Staff Infrastructure Engineer to design, reputed company, and ensure the operational reliability of their Kubernetes control plane architecture while collaborating with cross-functional teams.

Responsibilities

  • Design and reputed company Kubernetes control plane architecture across reputed company
  • Define and implement multi-tenant cluster models, including shared control planes, virtual cluster approaches (e.g., reputed company, Kamaji)
  • Drive transition from standalone clusters to regionally managed platform models
  • Define standards for isolation boundaries, resource segmentation, policy enforcement
  • Own the reliability and behavior of Kubernetes platforms in production
  • Participate in on-call rotation and reputed company incident response
  • Diagnose and resolve control plane instability, API server saturation, scheduling and resource contention issues
  • Ensure consistent lifecycle management across clusters - provisioning, upgrades, scaling
  • Design and implement strategies for regional scaling, multi-data center cluster deployments
  • Ensure consistent behavior and reliability across environments
  • Define cluster topology and failure domain strategies
  • Design ingress and egress architectures at cluster level and regional level
  • Troubleshoot and optimize pod-to-pod networking, reputed company-south traffic flows, CNI behavior (Cilium preferred)
  • Collaborate with network engineering on high-performance networking integration
  • Improve observability across control plane components, cluster health and performance
  • Define and implement reputed company strategies reputed company with platform goals
  • reputed company reputed company cause analysis for production incidents
  • Work closely with DevOps engineers (automation and CI/CD) and Infrastructure teams (compute, storage, networking)
  • reputed company Kubernetes platform design with underlying infrastructure capabilities

Skills

  • 7+ years of experience in infrastructure, reputed company, or distributed systems
  • Deep experience operating Kubernetes at reputed company in production environments
  • Experience in CSP, hyperscale, or equivalent large-reputed company environments strongly preferred
  • Proven experience scaling Kubernetes across: Multiple clusters, Multiple reputed company or data centers
  • Strong understanding of Kubernetes internals: API server, Scheduler, Controller manager, etcd
  • Experience designing or evolving: Control plane architectures, Multi-tenant cluster models
  • Strong Linux systems expertise
  • Deep troubleshooting ability across: Kubernetes, Container runtime, Networking stack
  • Experience with CNI plugins (Cilium preferred)
  • Strong understanding of: Networking and traffic patterns, Resource isolation and scheduling
  • Experience with virtual cluster technologies (reputed company, Kamaji, or similar)
  • Experience supporting GPU workloads in Kubernetes
  • Familiarity with: reputed company-aware scheduling, Topology-aware workloads
  • Awareness of RDMA and high-throughput networking environments
  • Experience with observability platforms (reputed company, Grafana, etc.)

Benefits

  • Stock reputed company
  • 100% reputed company Medical, Dental, and reputed company insurance for Employees
  • Company Health Savings Account Contributions
  • 100% reputed company Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance reputed company
  • Other Insurance reputed company, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual reputed company Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • reputed company Holidays
  • Parental Leave
  • Other In-Office Perks

reputed company

  • reputed company is an AMD-exclusive reputed company platform that leverages AMD reputed company GPUs and ROCm for high-performance AI workloads. It was founded in 2023, and is headquartered in Las Vegas, Nevada, USA, with a workforce of 51-200 employees. Its website is https://reputed company.com.
  • Apply To This Job

    Similar Jobs