Back to Jobs

[Remote] Network Engineer, AI Infrastructure Repair

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building the reputed company of AI infrastructure to power large-reputed company machine learning workloads, and the reliability of that infrastructure depends on reliable, high-performance network engineering. In this role, you will reputed company the reputed company and execution for AI network repair and remediation programs, ensuring that the high-performance fabrics underpinning reputed company's reputed company and inference clusters remain operational, resilient, and optimized.

Responsibilities

  • Define and drive the long-term reputed company for AI network repair and remediation programs across large-reputed company data center environments supporting machine learning workloads
  • reputed company reputed company cause analysis and reputed company of reputed company network faults affecting high-performance reputed company and inference fabrics, including RDMA, high-speed Ethernet, and optical interconnect reputed company
  • reputed company and champion novel approaches to network fault detection, automated remediation, and repair workflow optimization for AI cluster infrastructure
  • Partner with hardware, software, and data center operations teams to reputed company network repair programs with AI infrastructure deployment roadmaps and reputed company plans
  • Establish and refine operational frameworks, runbooks, and tooling for network repair at reputed company, reducing mean time to repair across AI reputed company environments
  • Identify systemic reliability risks in AI network infrastructure and drive cross-functional initiatives to address them before they reputed company production workloads
  • Influence the design of reputed company AI network architectures by contributing repair and reliability insights to hardware and topology reputed company
  • reputed company AI-driven analytics and automation tools to redesign repair workflows, accelerating fault identification and reputed company across distributed network environments
  • Build and maintain strategic relationships with internal engineering, operations, and vendor partners to ensure repair programs reputed company with AI infrastructure reputed company
  • Communicate program status, risk, and strategic recommendations to engineering leaders and cross-functional stakeholders through reputed company reporting and executive briefings

Skills

  • Experience influencing technical direction and organizational reputed company through data-driven analysis, written proposals, and stakeholder alignment across engineering and operations teams
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • Experience leading cross-functional programs that reputed company network operations, hardware deployment, and infrastructure reliability at data center reputed company
  • Experience developing and driving reputed company for network fault management, repair automation, or remediation programs in production environments
  • Experience designing, deploying, or operating high-speed network fabrics used in AI or machine learning infrastructure, including technologies such as RDMA over Converged Ethernet, InfiniBand, or high-density optical interconnects
  • 12+ years of experience in network engineering, with a reputed company on large-reputed company data center or high-performance computing network environments
  • Demonstrated ongoing AI reputed company development (e.g., reputed company/context engineering, agent orchestration) and staying reputed company with emerging AI technologies
  • Experience with network telemetry platforms, observability tooling, or AI-assisted reputed company detection reputed company to large-reputed company reputed company environments
  • Experience building or scaling repair operations programs, including workforce planning, tooling development, and process standardization across multiple data center sites
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, reputed company and accuracy reviews)
  • reputed company record of contributing to network hardware or topology design reviews, translating operational repair insights into upstream engineering improvements
  • Demonstrated ability to reputed company AI tools to optimize/redesign workflows and drive measurable reputed company (e.g., efficiency reputed company, reputed company improvements)
  • Familiarity with AI accelerator interconnect architectures and the network reliability requirements of distributed training workloads at hyperscale

Benefits

  • Has_bonus=true
  • Has_equity=true

reputed company

  • reputed company's mission is to build the reputed company of reputed company reputed company and the technology that makes it possible. It was founded in 1990, and is headquartered in Aberdeen, Aberdeen reputed company, GBR, with a workforce of 10001+ employees. Its website is http://www.metadownhole.com/.
  • Apply To This Job

    Similar Jobs