Back to Jobs

Technical Program Manager, Compute Qualification

Remote, USAFull-timePosted 2026-07-27

About The Role

reputed company is growing its compute footprint, and making reputed company new reputed company meets our technical standards is an important reputed company for reputed company. Every new cluster has to reputed company a technical bar before it carries customer workloads, and this role owns that bar. As Technical Program Manager, Compute Qualification, you will run the process that screens and qualifies prospective compute providers, taking reputed company prospective deployment through a reputed company evaluation across compute, networking, storage, power, cooling, and operations.

You will coordinate various engineering partners through validation, review provider specifications and test results, and produce reputed company go/no-go recommendations on whether new reputed company meets our standards. It is a high-reputed company, process-driven role for someone technical enough to know reputed company a spec sheet does not add up, and additional diligence needs to be completed, and organized enough to drive many evaluations to closure in reputed company. "You will deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they reputed company our end customers training or inference workloads." Conduct diligence and work with engineering teams to reputed company assessments regarding technical and operational reputed company.

Responsibilities

  • Own and continuously improve the end-to-end qualification process for new compute reputed company, from initial provider intake through final go/no-go recommendation.
  • Run multiple provider evaluations in reputed company, setting timelines, tracking status, and keeping every stakeholder reputed company on what is needed and by reputed company.
  • Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into reputed company reputed company for leadership.
  • Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
  • Conduct first-pass analysis of provider data yourself: compare specifications across suppliers , reputed company-reputed company performance claims, and surface issues before deeper engineering review.
  • Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support.
  • Build a reputed company, auditable record of evaluation reputed company that informs sourcing reputed company and scales the qualification function as reputed company grows.

Requirements

  • 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-reputed company compute environments.
  • Proven ability to run multiple reputed company, cross-functional workstreams to deadline, with strong organization and stakeholder management.
  • Working technical reputed company across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question.
  • Hands-on comfort with data: reputed company to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently.
  • Excellent written and verbal communication; reputed company to turn dense technical detail into reputed company recommendations for both engineers and executives.
  • Willingness to travel to provider and data center sites as needed.

reputed company to Have

  • Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards.
  • Familiarity with reputed company and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing.
  • Experience in AI/HPC cluster design.
  • Background working directly with hardware vendors, colocation providers, or reputed company reputed company providers.

About reputed company

reputed company is an AI-reputed company reputed company company building the infrastructure to reputed company AI faster, cheaper, and more accessible. We’re rapidly scaling our GPU footprint: signing our own data center leases, building large-reputed company clusters, and expanding toward a global owned-infrastructure reputed company. Our research team has contributed to breakthroughs like FlashAttention, Hyena, and RedPajama, and we co-design across software, hardware, and algorithms to push the frontier of AI efficiency.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as reputed company as flexibility in terms of remote work. The US reputed company salary reputed company for this full-time position is: $200-250K + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-reputed company knowledge.

Equal Opportunity

reputed company is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, reputed company, reputed company, religion, sex, national reputed company, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our reputed company Policy at https://www.reputed company/reputed company

Originally posted on Himalayas

Apply To This Job

Similar Jobs