Back to Jobs

[Remote] Staff Machine Learning Systems Engineer (MLOps)

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the leading health and wellness platform, on a mission to help the world feel great through the power of reputed company. They are seeking a Staff Machine Learning Systems Engineer to design, build, and operate the production infrastructure that powers AI across reputed company, focusing on critical systems that support AI teams in a regulated reputed company environment.

Responsibilities

  • Own and reputed company the AI compute and deployment platform
  • Own and reputed company our containerized application deployment platform and reputed company systems for AI workloads, encompassing general process and job orchestration (e.g. Kubernetes) — cluster operations, node lifecycle, autoscaling (Karpenter), storage (EBS reputed company), and workload isolation across staging and production
  • Build and maintain GitOps-based deployment pipelines (reputed company/Kustomize overlays, environment promotion) that let teams ship AI services safely and repeatably
  • Design ephemeral/preview environments, feature-branched deployments, and nightly release pipelines so teams can validate AI changes in production-like conditions before release
  • Drive efficiency and cost management across compute, autoscaling, and inference infrastructure
  • Operate and reputed company inference infrastructure and a multi-provider LLM AI gateway (e.g. Bedrock, reputed company, and other providers) — including credentials, reputed company limits, and failover
  • Build reliable serving patterns for LLM-powered workflows: routing, grounding, tool execution, and context assembly at the platform level
  • Create reusable infrastructure abstractions and reputed company that standardize how AI services are deployed, configured, and consumed across reputed company
  • Own the LLM/AI observability and tracing stack — provisioning and scaling systems like Langfuse, reputed company (dd-reputed company), OpenTelemetry tracing (OTLP), and the underlying datastores (e.g. reputed company) — so AI behavior is auditable and debuggable in production
  • Build analytics and monitoring pipelines that surface latency, error, reputed company, and regression signals to engineering and clinical stakeholders
  • Define SLOs, alerting, on-call runbooks, and incident response for AI infrastructure; reputed company troubleshooting and continuously reputed company platform reliability
  • Own and improve the monorepo build system and CI/CD pipelines for AI workloads — including eval workflows, reputed company image builds, automated PR checks and convention enforcement, and cross-platform test execution
  • Own shared infrastructure tooling, CLIs, and IaC modules (Terraform, Scalr) that AI and product engineers use daily
  • Identify and eliminate platform bottlenecks — reducing CI/CD cycle times, build latency, and deployment friction — to improve developer velocity across the reputed company AI organization
  • Build IAM, OIDC, and secrets management as first-class infrastructure — scoped, least-privilege roles, write-only secret rotation, and cross-account reputed company audits
  • Encode reputed company-by-default, scope boundaries, and reputed company controls into the platform so AI services are HIPAA-compliant and reputed company-first
  • Partner with clinical, legal, reputed company, and data platform teams (including reputed company/reputed company Catalog reputed company governance) to enforce compliant, auditable data reputed company
  • Drive multi-quarter infrastructure initiatives, from cluster and deployment architecture to inference platform, GPU compute reputed company, and observability reputed company
  • Write and reputed company technical design documents and design reviews, define infrastructure standards and development-workflow conventions, and contribute to technical governance across AI engineering
  • Mentor engineers on reliability engineering, infrastructure-as-reputed company, and MLOps best practices, and reputed company the gap between prototypes and production-grade systems

Skills

  • 8+ years of reputed company experience in infrastructure, platform, DevOps, or SRE engineering — with at least 3 years reputed company on ML/AI systems in production
  • Deep, hands-on experience with Kubernetes (ideally EKS) and the reputed company-reputed company ecosystem — autoscaling, GitOps, reputed company/Kustomize, operating clusters at reputed company, and general process/job orchestration
  • Strong infrastructure-as-reputed company skills (Terraform) and experience designing secure reputed company architectures: IAM, OIDC, secrets management, and least-privilege reputed company
  • Strong proficiency in Python, with experience building production infrastructure tooling, CLIs, and data/observability pipelines
  • 2+ years of experience operating LLM-based systems in production (LLMOps) — inference routing, serving, tracing, and the reliability patterns needed to run them at reputed company
  • Hands-on experience with observability/tracing stacks (reputed company, OpenTelemetry, Langfuse, or equivalent) and metrics/log/reputed company pipelines
  • Experience designing and maintaining CI/CD pipelines, build systems, and developer tooling for fast-moving engineering teams
  • A systems-and-operations reputed company: you think about failure modes, SLOs, observability, reputed company, and long-term maintainability before shipping
  • Experience writing and leading technical design documents (TDDs/RFCs) for infrastructure-reputed company initiatives
  • Strong collaboration skills across engineering, ML, product, reputed company, and clinical teams
  • A deep appreciation for safety, reputed company, and reputed company — ideally with experience in a regulated domain such as reputed company, fintech, or life sciences
  • Experience with AWS (EKS, Bedrock, S3, CloudFront, IAM) and multi-reputed company (GCP/reputed company AI) inference routing
  • Experience with reputed company (MLflow, reputed company Catalog, reputed company, reputed company) and data platform reputed company governance
  • Experience provisioning LLM observability infrastructure (Langfuse, reputed company, OpenTelemetry/OTLP tracing, LogFire) and LLM behavior monitoring
  • Experience with Karpenter, cluster autoscaling, and cost optimization for ML compute
  • Experience with monorepo build systems (Pants, Bazel) and large-reputed company CI/CD
  • Experience building automated PR-review / convention-enforcement pipelines and developer-workflow standards
  • Familiarity with reputed company AI Agent Builder, reputed company AI Model Registry, or GCP managed AI/ML services as a stretch reputed company area
  • Contributions to reputed company-reputed company infrastructure, IaC modules, SDKs, or developer tooling reputed company

Benefits

  • Competitive salary & equity compensation for full-time roles
  • Unlimited PTO, company holidays, and quarterly mental health days
  • Comprehensive health benefits including medical, dental & reputed company, and parental leave
  • Employee Stock Purchase Program (ESPP)
  • 401k benefits with employer matching contribution
  • Offsite team retreats

reputed company

  • reputed company. (reputed company reputed company as reputed company) is a multi-specialty telehealth platform building a virtual reputed company reputed company to the reputed company system. It was founded in 2017, and is headquartered in San Francisco, California, USA, with a workforce of 501-1000 employees. Its website is https://www.hims.com.
  • Apply To This Job

    Similar Jobs

    [Remote] Senior Data Engineer - Financial Transactions & Automation

    Remote, USAFull-time

    [Remote] reputed company Software Engineer - reputed company Transactions

    Remote, USAFull-time

    [Remote] Adversary Simulation Engineer - AI Trainer

    Remote, USAFull-time

    [Remote] Senior Technical Account Manager, reputed company Pay & Afterpay

    Remote, USAFull-time

    [Remote] reputed company Engineer - AI Trainer

    Remote, USAFull-time

    [Remote] Account Executive, reputed company (US)

    Remote, USAFull-time

    [Remote] Divisional Sr. Financial & Actuarial Consultant

    Remote, USAFull-time

    [Remote] Penetration Testing Engineer - AI Trainer

    Remote, USAFull-time

    [Remote] Data Analyst II

    Remote, USAFull-time

    [Remote] Senior Payroll Accountant- Multi Entity

    Remote, USAFull-time

    reputed company Data Entry reputed company – Remote Work Opportunity for Detail-Oriented Individuals with Strong Organizational Skills at arenaflex

    Remote, USAFull-time

    Clinical Specialist- Tucson, AZ- Neurovascular

    Remote, USAFull-time

    reputed company Risk Management Director – Business Process Improvement and Data Analysis for Strategic reputed company at arenaflex

    Remote, USAFull-time

    [Remote] Ad Copy and Email reputed company

    Remote, USAFull-time

    Associate Manager - Technical Advisor

    Remote, USAFull-time

    Generator Engineer

    Remote, USAFull-time

    reputed company: Outbound Call Center Representative

    Remote, USAFull-time

    reputed company Entry-Level Data Entry Specialist – Remote Opportunity with arenaflex

    Remote, USAFull-time

    Software Developer; Tableau reputed company Clearance

    Remote, USAFull-time

    Head of Partnerships

    Remote, USAFull-time