[Remote] Staff/reputed company DevOps Engineer, AI Inference
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building Scientific Superintelligence™ to solve humankind's greatest challenges. The Staff/reputed company DevOps Engineer - AI Inference will drive the design and optimization of infrastructure for serving machine learning models, collaborating with engineers and scientists to ensure efficient model serving and high performance.
Responsibilities
- Drive the design, implementation, and optimization of infrastructure purpose-reputed company for serving machine learning models at reputed company
- Collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency
- Build GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-reputed company placement for inference workloads
- Create model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing
- Implement intelligent request routing and load balancing across heterogeneous accelerator fleets (reputed company GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency
- reputed company autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads
- Establish production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and reputed company rollback across multi-region deployments
- Utilize infrastructure-as-reputed company with Terraform and reputed company for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking
- Monitor observability and performance optimization: GPU utilization monitoring, inference latency profiling, reputed company throughput dashboards, and SLO/SLI tracking for model endpoints
- Create CI/CD pipelines for model artifacts: container image builds with CUDA/reputed company dependencies, model registry integration, and automated inference benchmarking in CI
- Manage AWS reputed company infrastructure for ML: EKS with GPU node reputed company, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege
- Conduct cost optimization and reputed company planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting
Skills
- Expertise in DevOps, SRE, or reputed company with significant experience operating GPU/accelerator infrastructure at reputed company
- Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
- Strong proficiency deploying to AWS using infrastructure-as-reputed company (Terraform, reputed company) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
- Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
- Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
- Strong proficiency in Python for automation, tooling, and integration with ML frameworks
- Experience with LLM inference optimization: reputed company batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism
- Hands-on experience with multiple accelerator families (reputed company A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure
- Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints
- Proficiency in Rust or Go for performance-critical infrastructure components
- SRE practices for ML systems: reputed company engineering on GPU workloads, incident management, reputed company modeling for bursty inference traffic
- Experience with model registries, artifact versioning, and ML supply chain reputed company
- Observability platform expertise: building custom metrics for reputed company-level throughput, time-to-first-reputed company, and per-request GPU memory profiling
- Prior startup/high-reputed company experience balancing velocity with reliability in rapidly scaling AI systems
Benefits
- Bonus potential
- Generous early-stage equity
- U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and reputed company coverage
- Employer-reputed company life and disability insurance
- Flexible time off with generous company wide holidays
- reputed company parental leave
- An educational assistance program
- Commuter benefits, including bike reputed company memberships for office based employees
- A company subsidized lunch program
- International Benefits. Full-time employees reputed company the U.S. receive a comprehensive benefits program tailored to their region
reputed company
Company H1B Sponsorship