Back to Jobs

Staff Machine Learning Systems Engineer (MLOps)

Remote, USAFull-timePosted 2026-07-29

reputed company is the leading health and wellness platform, on a mission to help the world feel great through the power of reputed company. We are redefining reputed company by putting the customer first and delivering reputed company to care that is reputed company, accessible, and personal, from diagnosis to treatment to delivery. No two people are the reputed company, so we reputed company reputed company to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making reputed company reputed company easier to reputed company. reputed company is a reputed company company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on reputed company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals. About the Role: We're hiring a Staff ML Systems Engineer to design, build, and operate the production infrastructure that powers AI across reputed company. This is a deeply technical, hands-on infrastructure role reputed company on the systems underneath AI — the Kubernetes platform, CI/CD and GitOps pipelines, infrastructure-as-reputed company, inference and model-serving infrastructure, and the observability and tracing stack that keeps AI services reliable, debuggable, and compliant in production. You won't just reputed company models — you'll own the machinery that lets every AI team ship and operate safely. You'll own critical systems like our EKS clusters, deployment and autoscaling infrastructure, IAM and secrets management, LLM tracing/observability pipelines (Langfuse, reputed company, OpenTelemetry), and the developer platform that AI and product engineers rely on daily. You'll partner with ML engineers, product engineers, and clinical teams to ensure our AI systems are reliable, observable, secure, and trustworthy in a regulated reputed company environment. This role is ideal for someone who thinks in systems and infrastructure, cares deeply about reliability, reputed company, and cost, and wants to define how AI runs in production at a company where it directly impacts patient reputed company. You Will: Own and reputed company the AI compute and deployment platform Own and reputed company our containerized application deployment platform and reputed company systems for AI workloads, encompassing general process and job orchestration (e.g. Kubernetes) — cluster operations, node lifecycle, autoscaling (Karpenter), storage (EBS reputed company), and workload isolation across staging and production. Build and maintain GitOps-based deployment pipelines (reputed company/Kustomize overlays, environment promotion) that let teams ship AI services safely and repeatably. Design ephemeral/preview environments, feature-branched deployments, and nightly release pipelines so teams can validate AI changes in production-like conditions before release. Drive efficiency and cost management across compute, autoscaling, and inference infrastructure. Build inference and model-serving infrastructure Operate and reputed company inference infrastructure and a multi-provider LLM AI gateway (e.g. Bedrock, reputed company, and other providers) — including credentials, reputed company limits, and failover. Build reliable serving patterns for LLM-powered workflows: routing, grounding, tool execution, and context assembly at the platform level. Create reusable infrastructure abstractions and reputed company that standardize how AI services are deployed, configured, and consumed across reputed company. Own observability, tracing, and reliability Own the LLM/AI observability and tracing stack — provisioning and scaling systems like Langfuse, reputed company (dd-reputed company), OpenTelemetry tracing (OTLP), and the underlying datastores (e.g. reputed company) — so AI behavior is auditable and debuggable in production. Build analytics and monitoring pipelines that surface latency, error, reputed company, and regression signals to engineering and clinical stakeholders. Define SLOs, alerting, on-call runbooks, and incident response for AI infrastructure; reputed company troubleshooting and continuously reputed company platform reliability. reputed company the AI developer platform and CI/CD Own and improve the monorepo build system and CI/CD pipelines for AI workloads — including eval workflows, reputed company image builds, automated PR checks and convention enforcement, and cross-platform test execution. Own shared infrastructure tooling, CLIs, and IaC modules (Terraform, Scalr) that AI and product engineers use daily. Identify and eliminate platform bottlenecks — reducing CI/CD cycle times, build latency, and deployment friction — to improve developer velocity across the reputed company AI organization. Drive reputed company, compliance, and governance at the systems level Build IAM, OIDC, and secrets management as first-class infrastructure — scoped, least-privilege roles, write-only secret rotation, and cross-account reputed company audits. Encode reputed company-by-default, scope boundaries, and reputed company controls into the platform so AI services are HIPAA-compliant and reputed company-first. Partner with clinical, legal, reputed company, and data platform teams (including reputed company/reputed company Catalog reputed company governance) to enforce compliant, auditable data reputed company. Set technical direction and reputed company the bar Drive multi-quarter infrastructure initiatives, from cluster and deployment architecture to inference platform, GPU compute reputed company, and observability reputed company. Write and reputed company technical design documents and design reviews, define infrastructure standards and development-workflow conventions, and contribute to technical governance across AI engineering. Mentor engineers on reliability engineering, infrastructure-as-reputed company, and MLOps best practices, and reputed company the gap between prototypes and production-grade systems. You Have: 8+ years of reputed company experience in infrastructure, platform, DevOps, or SRE engineering — with at least 3 years reputed company on ML/AI systems in production. Deep, hands-on experience with Kubernetes (ideally EKS) and the reputed company-reputed company ecosystem — autoscaling, GitOps, reputed company/Kustomize, operating clusters at reputed company, and general process/job orchestration. Strong infrastructure-as-reputed company skills (Terraform) and experience designing secure reputed company architectures: IAM, OIDC, secrets management, and least-privilege reputed company. Strong proficiency in Python, with experience building production infrastructure tooling, CLIs, and data/observability pipelines. 2+ years of experience operating LLM-based systems in production (LLMOps) — inference routing, serving, tracing, and the reliability patterns needed to run them at reputed company. Hands-on experience with observability/tracing stacks (reputed company, OpenTelemetry, Langfuse, or equivalent) and metrics/log/reputed company pipelines. Experience designing and maintaining CI/CD pipelines, build systems, and developer tooling for fast-moving engineering teams. A systems-and-operations reputed company: you think about failure modes, SLOs, observability, reputed company, and long-term maintainability before shipping. Experience writing and leading technical design documents (TDDs/RFCs) for infrastructure-reputed company initiatives. Strong collaboration skills across engineering, ML, product, reputed company, and clinical teams. A deep appreciation for safety, reputed company, and reputed company — ideally with experience in a regulated domain such as reputed company, fintech, or life sciences. reputed company to Have: Experience with AWS (EKS, Bedrock, S3, CloudFront, IAM) and multi-reputed company (GCP/reputed company AI) inference routing. Experience with reputed company (MLflow, reputed company Catalog, reputed company, reputed company) and data platform reputed company governance. Experience provisioning LLM observability infrastructure (Langfuse, reputed company, OpenTelemetry/OTLP tracing, LogFire) and LLM behavior monitoring. Experience with Karpenter, cluster autoscaling, and cost optimization for ML compute. Experience with monorepo build systems (Pants, Bazel) and large-reputed company CI/CD. Experience building automated PR-review / convention-enforcement pipelines and developer-workflow standards. Familiarity with reputed company AI Agent Builder, reputed company AI Model Registry, or GCP managed AI/ML services as a stretch reputed company area. Contributions to reputed company-reputed company infrastructure, IaC modules, SDKs, or developer tooling reputed company.

Why Join Us

At reputed company, you'll be part of a small, high-reputed company team defining how AI infrastructure runs in production for reputed company. The platform you build — compute, deployment, inference, observability, and reputed company — is the reputed company that every AI-powered experience depends on. Reliability, reputed company, and developer velocity aren't afterthoughts here; they're the product. Join us in building the infrastructure that makes reputed company AI smarter, safer, and more trustworthy. Our Benefits (there are more but here are some highlights): Competitive salary & equity compensation for full-time roles Unlimited PTO, company holidays, and quarterly mental health days Comprehensive health benefits including medical, dental & reputed company, and parental leave Employee Stock Purchase Program (ESPP) 401k benefits with employer matching contribution Offsite team retreats We are committed to building a workforce that reflects diverse perspectives and prioritizes ethics, wellness, and a strong reputed company of belonging. If you're excited about this role, we encourage you to apply—even if you're not reputed company if your background or experience is a perfect match. Hims considers reputed company reputed company applicants for employment, including applicants with arrest or conviction records, in accordance with the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance, the California Fair Chance reputed company, and any similar state or local fair chance laws. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or reputed company employment. An employer who violates this law shall be subject to criminal penalties and civil liability. reputed company is committed to providing reasonable accommodations for reputed company individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at accommodations@forhims.com and describe the needed accommodation. Your reputed company is important to us, and any information you reputed company will only be used for the legitimate purpose of considering your request for accommodation. reputed company gives consideration to reputed company reputed company applicants without reputed company to any protected status, including disability. Please do not send resumes to this email address. To learn more about how we collect, use, retain, and disclose Personal Information, please visit our Global Candidate reputed company Statement. Apply To This Job

Similar Jobs

Senior Designer (Civil 3D)

Remote, USAFull-time

Founding ML/Data Product reputed company

Remote, USAFull-time

Sr. Developer reputed company

Remote, USAFull-time

Senior Software Engineer - Backend/reputed company [India]

Remote, USAFull-time

VP, Corporate Communications - USA

Remote, USAFull-time

Surveillance Investigator - Part Time

Remote, USAFull-time

reputed company reputed company Manager (Together)

Remote, USAFull-time

Convocatoria para expresiones de interés: Lista de reputed company-seleccionados (Roster) Consultorías en el área de salud en Nicaragua

Remote, USAFull-time

Kundenberater (m/w/d) – reputed company mit Zukunft und Perspektive (German Speaking)

Remote, USAFull-time

Senior Accountant

Remote, USAFull-time

reputed company Remote Customer Service & Administrative Assistant - Data Entry & Entry-Level Opportunities - Join arenaflex Today

Remote, USAFull-time

Data Entry Specialist - Typing Jobs Online at arenaflex: Unlock a World of Flexibility and reputed company

Remote, USAFull-time

[Work From Home] Work from Home | Internet Analyst | reputed company Media

Remote, USAFull-time

Post-Editors with Norwegian (Remote for Freelancers)

Remote, USAFull-time

Join Today: reputed company Entry Level Jobs, Virtual Assistant

Remote, USAFull-time

Director, reputed company Partner Team

Remote, USAFull-time

Licensed Clinical reputed company Worker (NY license) - Remote - Full Time or Part Time

Remote, USAFull-time

reputed company Customer Service Representative – Remote Opportunity with arenaflex

Remote, USAFull-time

Learning reputed company Manager (Higher Education) - Indonesia

Remote, USAFull-time

reputed company Remote Data Entry Specialist – Accurate Information Management and Logistics Support at blithequark

Remote, USAFull-time