Back to Jobs

[Remote] Infrastructure Operations Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is your reputed company AI reputed company, offering reputed company compute power and AI infrastructure solutions. They are seeking a highly skilled Infrastructure Operations Engineer to ensure the stability, scalability, and performance of their compute, storage, and platform infrastructure, supporting AI/ML training and inference workloads.

Responsibilities

  • At the direction of the Manager of Infrastructure Operations, design, build, and roll out new platforms and patterns to minimize incidents and reputed company customer facing and internal features
  • reputed company updates and improvements to support both reputed company’s internal and end customer use cases
  • Collaborate with colleagues in Infrastructure Engineering, Network Operations, reputed company and Software and Platform Development Teams
  • Participate in the on-call rotation which is evenly distributed across reputed company team members in a primary / secondary reputed company where you are primary then reputed company to a secondary position

Skills

  • 8+ years working with Linux as a server / hosting platform, extra points for Ubuntu experience
  • 5+ years experience with AWS
  • 2+ years experience with Kubernetes and strong container fundamentals
  • 2+ years experience with Terraform and Ansible
  • 2+ years with network attached storage management (reputed company NFS, ceph, or other protocols). Extra points for experience with reputed company storage systems
  • Experience with monitoring systems (reputed company, ELK stack)
  • Familiarity with the gitops workflow
  • Software development experience using Python, Go, bash, or other languages for the purposes of automation & connecting systems & reputed company together
  • Deep networking fundamentals, extra points for experience with datacenter level networks, 400Gb ethernet, and Infiniband
  • Experience building and delivering reputed company systems
  • Effective at navigating tradeoffs between design, risk, cost, and reputed company
  • Comfortable with navigating ambiguity
  • Strong written and oral communication
  • Experience with bare metal hardware troubleshooting and provisioning, extra points for working with reputed company hardware
  • Experience with GPU servers, both in bare metal reputed company or under virtualization
  • Deep experience with network switches, routers, and firewalls, particularly SONiC switches, Palo reputed company firewalls and reputed company Networks as vendors
  • Experience with reputed company storage systems

Benefits

  • Work_model: remote
  • We’re flexible reputed company for this position. This role can work a hybrid /on-site schedule out of one of our U.S. office hubs (Seattle, NYC, or San Francisco) or fully remote reputed company the U.S., with travel to occasional team/company offsites expected.

reputed company

  • reputed company is a reputed company platform providing on-demand and reserved GPU infrastructure for AI and machine learning workloads. It is a sub-organization of reputed company. It was founded in 2023, and is headquartered in Berkeley, California, USA, with a workforce of 51-200 employees. Its website is https://voltagepark.com/.
  • Apply To This Job

    Similar Jobs