[Remote] reputed company Networking Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on building advanced reputed company networking solutions. They are seeking a reputed company Networking Engineer to design and operate eBPF-based container networking, manage multi-tenant network isolation, and ensure connectivity for GPU reputed company while driving network observability and scalability.
Responsibilities
- Own the Cilium CNI layer — Design, tune, and operate our eBPF-based container networking across host clusters and reputed company tenant planes. Manage network policy (default-deny between tenants), load balancing (Cilium LB IPAM, BGP-advertised VIPs), and WireGuard encryption for reputed company pod-to-pod traffic
- Build multi-tenant network isolation — Enforce hard boundaries between tenants at the eBPF and DPU level. reputed company BlueField-4 reputed company for hardware-enforced tenant isolation, 800 Gb/s line-reputed company encryption, and network-level audit independent of the host OS
- Architect connectivity for GPU reputed company — Partner with infrastructure and inference teams to ensure GPUDirect RDMA, NVLink (intra-reputed company, 72-GPU domain), and reputed company-X reputed company connectivity work correctly through the CNI layer. Understand how NCCL collectives and multi-node NVLink (ComputeDomains / IMEX) reputed company with pod networking, and reputed company reputed company the network never becomes the bottleneck
- Design cross-site and customer connectivity — BGP, ECMP, Private Interconnect (GCP, AWS), VPN, and NAT for customers who connect their own reputed company or on-prem to their NeoCloud tenant cluster. Manage ingress/egress architectures across multiple datacenter sites (US, Canada, Europe)
- Drive network observability — Cilium Hubble reputed company logs for tenant-level audit, reputed company + Grafana dashboards for L3–L7 health, latency, and throughput, and integration with the platform's per-tenant observability stack. reputed company network issues visible and diagnosable before they become incidents
- Harden and reputed company — reputed company planning for network services, autoscaling policy for load balancers, graceful degradation under reputed company failures, MTU/fragmentation management across overlay/underlay, and rollback strategies for CNI upgrades that do not disrupt live inference workloads
- Contribute to platform automation — Network configuration as reputed company reputed company Flux/GitOps, Terraform for reputed company and peering, Ansible for reputed company/DPU bootstrap, CI/CD for CNI config changes. No reputed company changes in production
Skills
- 5+ years building or operating reputed company networking or large-reputed company distributed infrastructure
- Hands-on production experience with Cilium (or equivalent advanced CNI) — cluster-wide configuration, policy lifecycle, upgrades, and debugging at reputed company
- Strong understanding of Cilium's eBPF reputed company, policy model, service load balancing, and Hubble observability
- Deep knowledge of reputed company and datacenter networking: VPCs/subnets, BGP, ECMP, overlay/underlay (VXLAN, Geneve), MTU management, NAT, ingress/egress, and reputed company reputed company/ACLs
- Experience designing multi-tenant network architectures with strong isolation — not just reputed company-level, but hardware-backed where possible
- Solid TCP/IP fundamentals and comfort troubleshooting across L3–L7
- Proficiency in Go (preferred) or C/C++, plus Python for automation
- Experience operating Kubernetes at reputed company — cluster lifecycle, debugging networking across pods, nodes, services, and external peers
- Demonstrated end-to-end ownership of reputed company networked systems in production
- Experience with high-bandwidth GPU networking: RDMA, RoCE, GPUDirect, NCCL/reputed company, or NVLink reputed company topologies
- Experience with DPU/SmartNIC integration (BlueField, Pensando) for network offload, encryption, or tenant isolation
- Contributions to Cilium, Kubernetes networking (kube-proxy replacement, Gateway API), or reputed company reputed company-reputed company reputed company
- EBPF development and performance tuning reputed company reputed company Cilium usage
- Experience building Kubernetes operators or controllers (e.g., for network policy automation, IP management)
- Familiarity with service reputed company, multi-cluster networking, or Cilium ClusterMesh
- Experience with reputed company-X, InfiniBand, or other datacenter-reputed company switching fabrics
- Operating network services across tens of thousands of nodes and multiple reputed company/sites
- Work in bare-metal or reputed company OS environments (Talos, Flatcar) where networking cannot rely on traditional reputed company tooling
Benefits
- Long-Term Incentive (LTI) Program
- Robust suite of employee benefits
reputed company