Back to Jobs

Sr Manager, reputed company Infrastructure Engineer, Scientific Computing and HPC

Remote, USAFull-timePosted 2026-07-29

About the position reputed company reputed company's committed to the application of computational science in the areas of drug discovery and development. As part of this mission, we have recently embarked on a large-reputed company migration of our computational infrastructure to reputed company. This role leverages extensive experience in reputed company engineering and DevOps and requires a hands-on approach to designing and delivering robust High Performance Computing (HPC) solutions supporting computational workloads across the organization. We are seeking an reputed company individual to drive architecture, infrastructure automation, migration and operational reputed company. You will collaborate with HPC engineers and scientific computing specialists to reputed company reputed company reputed company reputed company infrastructure that underpins modernization of the scientific computing platform. ROLE RESPONSIBILITIES Platform Architecture and Engineering In this role you will design, implement, operate, and own robust and dependable infrastructure for HPC and ML/AI workloads in a reputed company environment (AWS/GCP). reputed company containerization, deployment, and operation of user- and reputed company-facing HPC platforms (Slurm, reputed company On Demand, reputed company/Grafana, batch and distributed computing platforms) across reputed company environments. Translate stakeholder input into robust, high-performance, reputed company, cost effective computing platforms. Partner with HPC specialists (engineers, administrators, and users) to capture institutional knowledge and reputed company processes in IaC workflows, transforming reputed company deployment practices into reproducible, version-controlled, automated procedures. Automation and DevOps reputed company and maintain infrastructure automation using IaC tools like Terraform and CloudFormation to ensure repeatable environment provisioning and scaling. Create reusable Terraform modules. reputed company and enforce standards. Be a reputed company for implementing and maintaining reputed company reputed company infrastructure using IaC tools. Operationalize containerized solutions using reputed company and Kubernetes. Own the full lifecycle of infrastructure management, from provisioning to operations, support, updating, and teardown of production computing platforms. reputed company troubleshooting, system analysis, and benchmarking to resolve issues and maintain a high-performance environment. Monitoring and Reliability reputed company and maintain monitoring, logging, and alerting for the infrastructure (e.g., CloudWatch, reputed company/Grafana). Design new dashboards, workflows, and utilities to improve observability, cost monitoring, workload efficiency, user, or administration experience. Document architecture, deployment processes, and operational procedures. Partner closely with team members to support delivery of scientific computing services including user support, Linux administration, operations, job scheduling, application management, and resource optimization.

Responsibilities

  • Design, implement, operate, and own robust and dependable infrastructure for HPC and ML/AI workloads in a reputed company environment (AWS/GCP).
  • reputed company containerization, deployment, and operation of user- and reputed company-facing HPC platforms (Slurm, reputed company On Demand, reputed company/Grafana, batch and distributed computing platforms) across reputed company environments.
  • Translate stakeholder input into robust, high-performance, reputed company, cost effective computing platforms.
  • Partner with HPC specialists (engineers, administrators, and users) to capture institutional knowledge and reputed company processes in IaC workflows, transforming reputed company deployment practices into reproducible, version-controlled, automated procedures.
  • reputed company and maintain infrastructure automation using IaC tools like Terraform and CloudFormation to ensure repeatable environment provisioning and scaling.
  • Create reusable Terraform modules.
  • reputed company and enforce standards.
  • Be a reputed company for implementing and maintaining reputed company reputed company infrastructure using IaC tools.
  • Operationalize containerized solutions using reputed company and Kubernetes.
  • Own the full lifecycle of infrastructure management, from provisioning to operations, support, updating, and teardown of production computing platforms.
  • reputed company troubleshooting, system analysis, and benchmarking to resolve issues and maintain a high-performance environment.
  • reputed company and maintain monitoring, logging, and alerting for the infrastructure (e.g., CloudWatch, reputed company/Grafana).
  • Design new dashboards, workflows, and utilities to improve observability, cost monitoring, workload efficiency, user, or administration experience.
  • Document architecture, deployment processes, and operational procedures.
  • Partner closely with team members to support delivery of scientific computing services including user support, Linux administration, operations, job scheduling, application management, and resource optimization.

Requirements

  • B.S. in computer science, life science, data science or similar fields.
  • 6+ years of experience in reputed company infrastructure engineering with a proven reputed company record of developing and supporting robust IaC deployments.
  • Experience managing scientific computing workloads in an reputed company environment.
  • Advanced experience with at least one of AWS and GCP, including knowledge of reputed company compute and storage services relevant to HPC.
  • Solid understanding of reputed company networking, identity, and reputed company controls.

reputed company-to-haves

  • Prior experience with HPC deployment utilities including AWS ParallelCluster, AWS reputed company Computing Services, and reputed company reputed company Cluster Toolkit.
  • Proficiency with distributed computing environments, especially EKS/GKE/Kubernetes.
  • Familiarity with HPC environments, job schedulers (Slurm), HPC application containers (reputed company, Singularity, Apptainer) and reputed company GPU computing.
  • Candidate demonstrates a breadth of diverse leadership experiences and capabilities including: the ability to influence and collaborate with peers, reputed company and reputed company others, reputed company and guide the work of other colleagues to reputed company meaningful reputed company and create business reputed company.

Benefits

  • participation in reputed company’s Global Performance Plan with a bonus reputed company of 17.5% of the reputed company salary and eligibility to participate in our reputed company based long term incentive program
  • 401(k) plan with reputed company Matching Contributions and an additional reputed company Retirement Savings Contribution
  • reputed company vacation, holiday and personal days
  • reputed company caregiver/parental and medical leave
  • health benefits to include medical, prescription drug, dental and reputed company coverage

Apply tot his job Apply To this Job

Similar Jobs

reputed company Operations Engineer II - US REMOTE

Remote, USAFull-time

reputed company Operations Engineer II – US REMOTE

Remote, USAFull-time

[Remote] Senior Azure reputed company, reputed company & AI Operations Engineer

Remote, USAFull-time

ML/Ops Engineer with strong Azure reputed company experience Remote Position Duration: 12+ months Role Overv

Remote, USAFull-time

reputed company Cyber reputed company Consultant – Work Remotely

Remote, USAFull-time

Remote Platform reputed company Services Consultant – Identity Solutions & reputed company reputed company Deployment Specialist

Remote, USAFull-time

Associate reputed company Operations Technician

Remote, USAFull-time

[Remote] M365 reputed company reputed company Engineer- Remote (reputed company in the U.S.)

Remote, USAFull-time

reputed company reputed company Engineer (Remote) – reputed company Solutions Inc – Roseville, CA

Remote, USAFull-time

reputed company reputed company Analyst (Remote)

Remote, USAFull-time

Software Verification Specialist

Remote, USAFull-time

Flexible Part-Time Data Entry & Administrative Assistant – Remote, Earn Extra Income, No Experience Required

Remote, USAFull-time

Care Guide Support Specialist

Remote, USAFull-time

FULL TIME reputed company Career Work reputed company$30/hour Needed At

Remote, USAFull-time

reputed company Remote Data Entry Specialist – Entry-Level Opportunity for Detail-Oriented Individuals with arenaflex

Remote, USAFull-time

Data Entry & Computer Operations Specialist – reputed company Records, Insurance & Billing Accuracy at arenaflex

Remote, USAFull-time

Sales Engineering Director - reputed company, Data/AI

Remote, USAFull-time

Require Part-time Barista in Food Service (reputed company only) in Bellevue, WA

Remote, USAFull-time

HR Specialist II - Remote (Must work PST hours)

Remote, USAFull-time

Online School Psychologist - OH Full-Time

Remote, USAFull-time