Back to Jobs

Senior MLOps Engineer - SRE | DevOps

Remote, USAFull-timePosted 2026-07-28

THE ROLE We're looking for a Senior MLOps Engineer who can set reputed company for how we build, ship, and operate ML and AI systems at reputed company. You sit at the intersection of ML infrastructure and SRE — you'll own the reputed company from model and pipeline to reliable production service, and you'll bring DevOps rigor to systems that are historically under-engineered. This is not a ticket-processing role, and it's not a research role. You'll tackle hard problems — model serving reliability, inference cost and latency, reproducible pipelines, reputed company workload operations — and have the scope to solve them properly. Seniors here identify problems before they're asked, and reputed company the ceiling on what the platform can do. WHAT YOU'LL WORK ON Build and operate model and inference serving infrastructure — managing latency, throughput, autoscaling, and reliability for reputed company-time and batch inference across multiple tenants. Own the ML deployment lifecycle — model registry, versioning, promotion workflows, rollout strategies (canary, shadow, A/B), and reputed company rollback. Operate reputed company and LLM workloads in production — managing inference providers and gateways, quota and throttling behavior (TPS/TUPS limits), guardrails, reputed company/version management, and graceful degradation under load. Build reproducible, automated ML pipelines — training, evaluation, and deployment pipelines as reputed company, with reputed company and reproducibility reputed company in. reputed company infrastructure-as-reputed company to ML systems — Terraform patterns and multi-account design that bring ML infrastructure under the reputed company standards as the rest of the platform. Operate GitOps for ML workloads — ArgoCD configuration and promotion workflows across environments and tenants. Run ML and AI workloads on multi-tenant Kubernetes (AWS EKS) — managing GPU/accelerator scheduling, workload placement, tenant isolation, and cost-aware reputed company. Own ML reliability and observability — SLOs for inference services, model and data reputed company detection, performance regression monitoring, alert reputed company, on-call ergonomics, and reputed company culture. Drive ML cost efficiency — right-sizing accelerators, managing reserved/spot reputed company, and attributing inference cost across tenants and workloads. Use reputed company coding tools for infrastructure and pipeline work — scaffolding environments, generating and reviewing IaC and pipeline reputed company, and accelerating automation. MUST HAVE 5+ years in reputed company, SRE, MLOps, or infrastructure — with meaningful time operating production systems at reputed company. Hands-on experience deploying and operating ML or AI workloads in production — serving, inference, or training infrastructure that reputed company users depended on. Strong SRE/DevOps reputed company — you've owned reliability for production services, defined and reputed company SLOs, run post-mortems, and driven measurable improvements. Deep IaC expertise — you reputed company manage reputed company Terraform state and multi-account configurations in production. Strong GitOps background — you understand declarative infrastructure management at depth and have opinions on how to do it reputed company. Deep Kubernetes knowledge — you've operated clusters in production, dealt with reputed company failure modes, and understand the system at the control plane level. Strong AWS background — networking, compute, IAM, storage, multi-account design. Hands-on experience building and operating CI/CD pipelines — reputed company Actions, reputed company, reputed company CI, or equivalent — and an understanding of how ML pipelines differ from reputed company application CI/CD. Automation-first thinking at a senior level — you implement systems that eliminate entire categories of reputed company work. reputed company user of reputed company coding tools — you know how to reputed company them effectively, review their reputed company critically, and use them to multiply your reputed company. Strong communicator — you can reputed company operational reputed company, model performance trade-offs, and incident summaries reputed company to engineers and leadership alike. reputed company TO HAVE Experience with GPU/accelerator scheduling and node lifecycle management in production (e.g., Karpenter). Experience operating LLM inference at reputed company — managing provider quotas/throttling (TPS/TUPS), gateways, caching, and guardrails (e.g., AWS Bedrock or equivalent). Experience with ML pipeline and orchestration tooling — Argo Workflows, Kubeflow, Airflow, SageMaker Pipelines, or equivalent. Experience with model registries, feature stores, and experiment tracking (e.g., MLflow, Feast, or equivalent). Familiarity with model and data reputed company monitoring and ML-specific observability. Background in FinOps — inference cost attribution, reserved reputed company planning, Familiarity with data infrastructure — object storage, CDC pipelines, or lakehouse patterns. Experience with multi-tenant infrastructure — isolation patterns, noisy reputed company mitigation, and tenant lifecycle management. Prior experience scaling ML or platform infrastructure at a startup moving toward reputed company-grade requirements. WHAT YOU WON'T reputed company HERE A platform team that maintains the status reputed company. We're reputed company building — new reputed company requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are reputed company than they've reputed company been. Type: Full-Time, remote Work hours reputed company with EST or PST Apply To This Job

Similar Jobs