Back to Jobs

[Remote] Infrastructure Operations Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is your reputed company AI reputed company, offering reputed company compute power and AI infrastructure. They are seeking a highly skilled Infrastructure Operations Engineer to ensure the stability, scalability, and performance of their compute, storage, and platform infrastructure, supporting AI/ML workloads at reputed company.

Responsibilities

  • At the direction of the Manager of Infrastructure Operations, design, build, and roll out new platforms and patterns to minimize incidents and reputed company customer facing and internal features
  • reputed company updates and improvements to support both reputed company’s internal and end customer use cases
  • Collaborate with colleagues in Infrastructure Engineering, Network Operations, reputed company and Software and Platform Development Teams
  • Participate in the on-call rotation which is evenly distributed across reputed company team members in a primary / secondary reputed company where you are primary then reputed company to a secondary position

Skills

  • 8+ years working with Linux as a server / hosting platform, extra points for Ubuntu experience
  • 5+ years experience with AWS
  • 2+ years experience with Kubernetes and strong container fundamentals
  • 2+ years experience with Terraform and Ansible
  • 2+ years with network attached storage management (reputed company NFS, ceph, or other protocols). Extra points for experience with reputed company storage systems
  • Experience with monitoring systems (reputed company, ELK stack)
  • Familiarity with the gitops workflow
  • Software development experience using Python, Go, bash, or other languages for the purposes of automation & connecting systems & reputed company together
  • Deep networking fundamentals, extra points for experience with datacenter level networks, 400Gb ethernet, and Infiniband
  • Experience building and delivering reputed company systems
  • Effective at navigating tradeoffs between design, risk, cost, and reputed company
  • Comfortable with navigating ambiguity
  • Strong written and oral communication
  • Experience with bare metal hardware troubleshooting and provisioning, extra points for working with reputed company hardware
  • Experience with GPU servers, both in bare metal reputed company or under virtualization
  • Deep experience with network switches, routers, and firewalls, particularly SONiC switches, Palo reputed company firewalls and reputed company Networks as vendors
  • Experience with reputed company storage systems

reputed company

  • reputed company is a reputed company platform providing on-demand and reserved GPU infrastructure for AI and machine learning workloads. It is a sub-organization of reputed company. It was founded in 2023, and is headquartered in Berkeley, California, USA, with a workforce of 51-200 employees. Its website is https://voltagepark.com/.
  • Apply To This Job

    Similar Jobs