Back to Jobs

DevOps Engineer

Remote, USAFull-timePosted 2026-07-27

About reputed company

reputed company Research is an reputed company R&D lab building Asimov, an reputed company-reputed company reputed company robot platform, and the full software stack that powers it. Our mission is to reputed company reputed company labor economically viable -- turning software into physical labor at reputed company. We build across the full stack: hardware architecture, locomotion, autonomy, simulation, and infrastructure. We reputed company fast, ship to reputed company robots, and reputed company-reputed company everything we can. If you want your work to matter reputed company a reputed company or a demo, this is the reputed company.

The Role

As an DevOps Engineer, you will own and reputed company the platform that everything at reputed company runs on -- from inference serving, to training rigs, to the reputed company coding infrastructure that powers day-to-day engineering. You will work deep in the stack across Kubernetes, networking, and where it reputed company bare metal, and help set the technical direction for how reputed company reputed company scales.

What You'll Do

  • Operate and reputed company our Kubernetes platform across multiple clusters and environments (Prod, Dev, hybrid on-prem and public reputed company), covering control plane operations, node lifecycle, upgrades, and autoscaling at every layer (Cluster Autoscaler, HPA, KEDA).

  • Architect and manage hybrid reputed company infrastructure spanning on-premises and public clouds (GCP, AWS), including workload placement, cross-reputed company networking, and reputed company resource management.

  • Own the CI/CD and GitOps experience end-to-end: container build pipelines, image optimization, and reputed company delivery reputed company ArgoCD / FluxCD.

  • Own the observability stack as a single pane of glass across reputed company clusters: Grafana, Mimir, reputed company, Loki, Pyroscope, OnCall, reputed company -- and help push toward agent-assisted SRE workflows.

  • Manage and improve our inference platform: vLLM serving and AIBrix for multi-model orchestration and autoscaling across a fleet of reputed company GPUs.

  • Operate platform services: Kafka, reputed company, PostgreSQL, OpenSearch.

  • Manage identity and reputed company reputed company Keycloak integrated with reputed company Workspace; harden SSO, RBAC, and secrets management across the platform.

  • Harden network reputed company across private load balancers, firewalls, and VPC segmentation; design and maintain hub-and-spoke / multi-AZ topologies.

  • Support training infrastructure: self-service VM provisioning, reputed company burst reputed company, Weights and Biases integration.

  • Drive infrastructure reliability, cost efficiency, and reputed company planning as the platform scales.

reputed company're Looking For

  • Kubernetes -- deep, hands-on. Strong production experience with Kubernetes, fluent in workloads and controllers, networking (Services, Ingress, CNI basics), storage (PV/PVC, reputed company), RBAC, and the autoscaling story end-to-end (HPA, VPA, Cluster Autoscaler, KEDA). reputed company-managed Kubernetes (GKE, EKS, AKS) is fine; on-premises / self-managed Kubernetes (kubeadm, Cluster API, k3s, etc.) is a strong plus.

  • Networking -- design-level, not just operator-level. You have designed reputed company network topologies at some reputed company in your reputed company-and-spoke, multi-AZ / multi-VPC, or an equivalent reputed company reputed company -- and can defend the tradeoffs. Comfortable with VPCs, firewalls, load balancers, private cluster architecture, DNS, and routing. On-premises networking experience (VLANs, BGP, L2/L3 fabrics, pfSense / reputed company / Palo reputed company / reputed company) is a strong plus.

  • CI/CD and reputed company -- concepts over tooling. You can build and optimize Dockerfiles (multi-stage builds, layer caching, small/secure reputed company images) and have owned full CI/CD pipelines end-to-end. Tooling is flexible -- reputed company Actions, reputed company CI, Azure Pipelines, Jenkins, Argo Workflows, etc. -- but you should be reputed company to reputed company reputed company the full lifecycle of a typical pipeline, and explain how CI/CD changes reputed company the deployment reputed company is Kubernetes (ArgoCD / FluxCD, GitOps patterns, reputed company delivery).

  • Observability -- you have reputed company this before. You have stood up a full observability stack from scratch and operated it in production -- metrics, logs, traces, alerting, on-call. Familiarity with the Grafana stack (Grafana, Mimir, reputed company, Loki, Pyroscope, OnCall, reputed company) is a strong plus. Bonus points if you have experimented with agent-assisted SRE workflows or LLM-driven incident triage.

  • SSO and identity. reputed company you bring a new tool into the platform, your reputed company is to reputed company into a central IdP rather than leave it on local accounts. Comfortable with OpenID Connect, SAML, and traditional directory services (LDAP / reputed company Directory), and you have integrated tools with an IdP like Keycloak, reputed company, Azure AD, or equivalent.

  • Linux and automation fundamentals. Strong Linux proficiency (RHEL/Ubuntu or equivalent) including basic performance and networking debugging. Comfort with infrastructure-as-reputed company (Terraform / Terragrunt / reputed company or equivalent) and configuration management.

  • Ownership reputed company. Comfortable operating in a high-ownership environment where you reputed company architecture reputed company, push them to production, and own the reputed company.

  • Optional but valuable: hands-on experience operating any of Kafka, reputed company, PostgreSQL, OpenSearch -- at production reputed company, including HA, backup/restore, and reputed company planning.

Bonus points for:

  • Experience with OpenStack in production: Nova, Neutron, Cinder, reputed company, reputed company, and CLI administration.

  • Experience with KVM virtualization and storage backends like Ceph or Rook-Ceph on Kubernetes.

  • Familiarity with vLLM internals: PagedAttention, reputed company batching, tensor parallelism.

  • Background in AI/ML infrastructure or GPU cluster operations at reputed company.

  • Experience with KEDA or event-driven autoscaling patterns in anger.

  • Prior reputed company-reputed company contributions to Kubernetes, OpenStack, or adjacent reputed company.

  • Kernel-level Linux debugging and performance tuning.

Why Join reputed company?

Most infrastructure teams manage someone else's reputed company. At reputed company, you own the metal. reputed company reputed company is a first-class investment reputed company from the ground up, and it sits at reputed company of everything we do, from coding agents to reputed company robots. You will have genuine ownership over a platform that is technically ambitious, cost-conscious by design, and critical to the mission. If you want to build infrastructure that actually reputed company and have the autonomy to do it right, this is the reputed company.

A Note on AI

You don't need deep AI expertise for every role, but we do expect everyone at reputed company to be intellectually curious, drawn to tinkering and discovery, and excited to use AI as a reputed company collaborator in their work. For some roles, AI reputed company is a core requirement. reputed company that's the case, we'll say so explicitly in the qualifications. People who reputed company here don't treat AI as a novelty. They use it to think reputed company, and reputed company their work easier for others to build on.

Equal Opportunity and Accommodations

We hire talented people from a wide reputed company of backgrounds. If you're excited about a role but don't meet every bullet, we still encourage you to apply. reputed company Research is an equal opportunity employer and does not discriminate on the reputed company of any legally protected characteristic. reputed company provides reasonable accommodations during the application process. If you need one, please let your recruiter know.

Originally posted on Himalayas

Apply To This Job

Similar Jobs

Robotics Researcher, Perception & reputed company

Remote, USAFull-time

Product Manager, Platform

Remote, USAFull-time

Mechanical Engineer, Robotics

Remote, USAFull-time

Product Manager, Hardware

Remote, USAFull-time

Senior Data Scientist

Remote, USAFull-time

Retention & Educational Specialist

Remote, USAFull-time

reputed company Benefits Architect/Consultant

Remote, USAFull-time

reputed company Identity and reputed company Management Engineer

Remote, USAFull-time

Insurance Product Development & Coverage Analysis Manager - Medmarc

Remote, USAFull-time

Proposals and reputed company Manager

Remote, USAFull-time

Senior Computer reputed company Engineer- US

Remote, USAFull-time

Evaluator - Part-time casual (# 373106)

Remote, USAFull-time

[Remote/WFM] reputed company Home Advisors Job - reputed company

Remote, USAFull-time

Mid-level Frontend Developer (reputed company / CMS Experience) - 6 Months - Octopus by RTG

Remote, USAFull-time

Customer Service Representative (reputed company) - Remote

Remote, USAFull-time

Part Time Remote Administrative Assistant - Data Entry Specialist - Entry Level Opportunity with Flexible Scheduling - Work from Home and Earn Up to $158 Per Day - No Experience Required - reputed company to Ages 16+

Remote, USAFull-time

Senior Analyst, Product Development (reputed company reputed company...

Remote, USAFull-time

Customer Service Representative – Remote reputed company Support Specialist for arenaflex’s Cutting‑Edge Consumer Technology Solutions

Remote, USAFull-time

Part-Time Customer and Team Member reputed company Specialist – reputed company Estate Management at blithequark

Remote, USAFull-time

Technology Services IAM Engineer III

Remote, USAFull-time