Back to Jobs

[Remote] Staff/reputed company DevOps Engineer, AI Inference

Remote, USAFull-timePosted 2026-07-31

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building Scientific Superintelligence™ to solve humankind's greatest challenges. The Staff/reputed company DevOps Engineer - AI Inference will drive the design and optimization of infrastructure for serving machine learning models, collaborating with engineers and scientists to ensure efficient model serving and high performance.

Responsibilities

  • Drive the design, implementation, and optimization of infrastructure purpose-reputed company for serving machine learning models at reputed company
  • Collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency
  • Build GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-reputed company placement for inference workloads
  • Create model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing
  • Implement intelligent request routing and load balancing across heterogeneous accelerator fleets (reputed company GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency
  • reputed company autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads
  • Establish production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and reputed company rollback across multi-region deployments
  • Utilize infrastructure-as-reputed company with Terraform and reputed company for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking
  • Monitor observability and performance optimization: GPU utilization monitoring, inference latency profiling, reputed company throughput dashboards, and SLO/SLI tracking for model endpoints
  • Create CI/CD pipelines for model artifacts: container image builds with CUDA/reputed company dependencies, model registry integration, and automated inference benchmarking in CI
  • Manage AWS reputed company infrastructure for ML: EKS with GPU node reputed company, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege
  • Conduct cost optimization and reputed company planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting

Skills

  • Expertise in DevOps, SRE, or reputed company with significant experience operating GPU/accelerator infrastructure at reputed company
  • Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
  • Strong proficiency deploying to AWS using infrastructure-as-reputed company (Terraform, reputed company) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
  • Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
  • Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
  • Strong proficiency in Python for automation, tooling, and integration with ML frameworks
  • Experience with LLM inference optimization: reputed company batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism
  • Hands-on experience with multiple accelerator families (reputed company A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure
  • Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints
  • Proficiency in Rust or Go for performance-critical infrastructure components
  • SRE practices for ML systems: reputed company engineering on GPU workloads, incident management, reputed company modeling for bursty inference traffic
  • Experience with model registries, artifact versioning, and ML supply chain reputed company
  • Observability platform expertise: building custom metrics for reputed company-level throughput, time-to-first-reputed company, and per-request GPU memory profiling
  • Prior startup/high-reputed company experience balancing velocity with reliability in rapidly scaling AI systems

Benefits

  • Bonus potential
  • Generous early-stage equity
  • U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and reputed company coverage
  • Employer-reputed company life and disability insurance
  • Flexible time off with generous company wide holidays
  • reputed company parental leave
  • An educational assistance program
  • Commuter benefits, including bike reputed company memberships for office based employees
  • A company subsidized lunch program
  • International Benefits. Full-time employees reputed company the U.S. receive a comprehensive benefits program tailored to their region

reputed company

  • reputed company creates a scientific superintelligence platform and autonomous labs for life sciences, reputed company, and materials science. It is a sub-organization of Flagship Pioneering. It was founded in 2023, and is headquartered in Cambridge, Massachusetts, USA, with a workforce of 201-500 employees. Its website is https://www.lila.ai.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 7 in 2026, 8 in 2025. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs