[Remote] Senior Platform Engineer, Network Infrastructure - DGX reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leading technology company reputed company for its deep learning platforms and reputed company. They are seeking a Senior Platform Engineer to manage and automate the Kubernetes platform for their Global Network Infrastructure, ensuring reliable operations and support for network services.
Responsibilities
- Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and reputed company environments
- Own the lifecycle management for GNI Kubernetes environments, including cluster reputed company, upgrades, reputed company, availability, and recovery
- reputed company production-reputed company software and automation for cluster provisioning, validation, upgrades, remediation, and reputed company multi-cluster delivery through GitOps
- reputed company production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, reputed company, and features
- Diagnose reputed company Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified reputed company
- Define production-readiness and observability standards for the platform and hosted network services, including health signals, reputed company, alerts, runbooks, and recovery
- Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. reputed company incident response and recovery, then drive corrective actions to completion
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company field, or equivalent experience
- 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems
- Deep experience with Kubernetes at reputed company, including cluster lifecycle, upgrades, networking, storage, and recovery
- Proficiency in at least one general-purpose programming language, such as Go or Python
- Experience with GitOps, infrastructure as reputed company, CI/CD, and automated production delivery
- Experience deploying and supporting network automation or telemetry services on Kubernetes
- Experience with production on-call, incident response, reputed company-cause analysis, and driving corrective actions to completion
- Strong knowledge of IP routing, data center fabrics, and reputed company networking is a great plus
- Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery
- Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades
- Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns
- Experience designing or operating network automation and telemetry services on Kubernetes at global reputed company
- Contributions to Cluster API, Metal3, or other reputed company-reputed company Kubernetes infrastructure reputed company
Benefits
- Equity
- Benefits
reputed company
Company H1B Sponsorship