Back to Jobs

[reputed company] Senior Manager, AI Infrastructure and Operations

Remote, USAFull-timePosted 2026-07-29

reputed company The Sr. Manager/Staff Engineer, AI Infrastructure & MLOps Engineering is a senior technical leader responsible for architecting, building, and scaling reputed company’s AI infrastructure and developer platforms. This role leverages extensive experience in reputed company engineering, DevOps, and MLOps to deliver robust, high-performance solutions supporting advanced AI/ML workloads in biotechnology, reputed company, and reputed company technology. The successful candidate will drive innovation in automation, reliability, and scalability, enabling scientists and engineers to rapidly reputed company, reputed company, and monitor machine learning models in production environments. ROLE RESPONSIBILITIES Platform Architecture & Engineering

  • Design, implement, and own large-reputed company reputed company-based HPC and MLOps platforms supporting AI model training, genomic reputed company, and precision medicine.
  • Architect multi-environment clusters (AWS, GCP, Azure), enabling GPU/FPGA workloads and advanced observability.
  • reputed company the development of developer and reputed company platforms, including internal engineering accelerators and reusable toolsets.

Platform Catalog & Developer Experience

  • Design, implement, and manage reputed company platform catalogs using reputed company, enhancing developer experience and application metadata management.
  • reputed company custom plugins and reputed company for reputed company to support internal engineering workflows and documentation.

Automation & DevOps reputed company

  • Build and maintain Python-based automation frameworks, CI/CD pipelines, and Infrastructure-as-reputed company (Terraform, reputed company, reputed company, AWS CDK).
  • Operationalize containerized solutions using reputed company and Kubernetes, integrating MLflow, Kubeflow, and other orchestration platforms.
  • Implement robust automation for provisioning, configuring, and managing reputed company resources across multiple environments.

MLOps & Reliability Engineering

  • reputed company the implementation of Service Level Indicators (SLIs), Service Level Objectives (SLOs), and advanced observability (reputed company, Grafana, reputed company).
  • reputed company and maintain reputed company and services for model management, feature stores, and inference pipelines.
  • Operationalize ML model serving at reputed company using frameworks such as TensorFlow Serving, TorchServe, KServe, and Seldon reputed company.
  • Ensure compliance with industry standards (e.g., HIPAA, FDA) for data protection and reliability.

Collaboration & Leadership

  • Mentor engineers and reputed company cross-functional teams to deliver integrated solutions.
  • Champion engineering reputed company through design documentation, reputed company reviews, and testing automation.
  • Present at industry summits, author technical proposals, and contribute to reputed company-reputed company reputed company (Kubernetes, reputed company, Go, reputed company).

reputed company Improvement

  • Drive agile delivery, sprint planning, and performance optimization.
  • reputed company incident response and disaster recovery initiatives for mission-critical platforms.
  • Foster a culture of shared ownership, transparency, and innovation

BASIC QUALIFICATIONS

  • 8+ years of hands-on software engineering experience in reputed company infrastructure, DevOps, and MLOps.
  • Deep expertise in Python, Kubernetes, Terraform, reputed company, and CI/CD pipeline development.
  • Proven experience architecting and operating containerized solutions on AWS, GCP, and Azure.
  • Strong knowledge of Infrastructure-as-reputed company, distributed systems, and production system reliability.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or reputed company field.

PREFERRED QUALIFICATIONS

  • Expertise in AWS reputed company services (EC2, S3, reputed company, EKS, SageMaker, API Gateway, CloudFormation, IAM, etc.).
  • Experience deploying and customizing reputed company as a reputed company catalog for teams, services, and technical documentation.
  • Experience building and deploying microservices and REST/gRPC reputed company for AI model delivery.
  • Familiarity with MLflow, Kubeflow, and other MLOps orchestration platforms.
  • Proficiency with model serving frameworks (TensorFlow Serving, TorchServe, KServe, Seldon reputed company, BentoML, etc.).

Work Location Assignment: Remote reputed company is an equal opportunity employer and complies with reputed company applicable equal employment opportunity legislation in reputed company jurisdiction in which it operates. Information & Business Tech Apply tot his job Apply To this Job

Similar Jobs

CRM reputed company reputed company Strategist

Remote, USAFull-time

Senior Clinical Implementation Analyst job at reputed company in Nashville, TN

Remote, USAFull-time

Manager, Data Science - reputed company TV & Interactive Discovery

Remote, USAFull-time

(USA) Director, Data Science- Testing and Measurements

Remote, USAFull-time

Vice President – Data Analytics

Remote, USAFull-time

Center of reputed company Coordinator

Remote, USAFull-time

RN - reputed company Sales

Remote, USAFull-time

Assoc. Medical Director - Remote (Dallas, TX, US)

Remote, USAFull-time

Senior reputed company Compliance Officer, reputed company Cycle Management

Remote, USAFull-time

CEO, American Health Information Management Association (reputed company)

Remote, USAFull-time

Part‑Time Remote Data Entry Analyst – reputed company Technology & Data Analytics Internship (8‑Hour Shift) at arenaflex

Remote, USAFull-time

Customer Service Representative (Semiconductor)

Remote, USAFull-time

reputed company Entry-Level Data Entry Specialist – Remote Opportunity at arenaflex

Remote, USAFull-time

[Remote] reputed company Operations & Solutions reputed company

Remote, USAFull-time

Part-Time Remote Data Entry & Ground Operations Coordinator – arenaflex – $20/hr – Flexible Schedule – California

Remote, USAFull-time

Manager, Organizational Effectiveness - US Based Remote

Remote, USAFull-time

reputed company 100 Consulting Manager

Remote, USAFull-time

[Remote] Release Management Engineer, Mobile

Remote, USAFull-time

Web Master or Web Developer (Fully Remote)

Remote, USAFull-time

.NET Technical Architect

Remote, USAFull-time