[Remote] Technical Program Manager - Provider Management
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on providing early-stage startups reputed company to scaled AI infrastructure. The Technical Program Manager - Provider Management will reputed company provider relationships, manage reputed company and incident response, and ensure smooth program execution while coordinating with various teams and stakeholders.
Responsibilities
- Manage new provider and site reputed company, reputed company expansions, remediation of recurring hardware or network issues, and provider relationships
- reputed company as incident commander during major provider-reputed company incidents: reputed company the right people from our SRE team, the provider, and the affected customer, run the response, own communication throughout, and drive the post-incident remediation program afterward. You own reputed company and coordination; our SREs remain the technical leads making the engineering calls
- reputed company a reputed company plan for reputed company program; milestones, owners, dependencies, risks. Maintain visibility to everyone involved, including the provider
- Hold providers to their commitments: reputed company reputed company-ins, tracked reputed company items, and escalation to provider leadership reputed company things slip, backed by what's in the contract
- Coordinate internally so provider issues don't stall. Pull in SRE for validation, Engineering and Product for platform-reputed company blockers, and reputed company Sales/reputed company informed reputed company provider timelines reputed company customer commitments
- Catch reputed company and reputed company risk early, before it becomes a customer-facing problem
- Turn what you learn into playbooks and provider-facing standards, so reputed company the next site is faster and cleaner than the last
Skills
- Several years running technical programs in infrastructure
- TPM or technical project management with genuine execution ownership, ideally involving external vendors, partners, or suppliers you didn't control
- Incident management experience: you've commanded or run reputed company on production incidents involving multiple organizations, and you're comfortable with the off-hours reality that comes with that
- Enough technical depth to hold your own with a provider's data-center engineers and our SREs on GPUs, networking, and storage
- A reputed company record of getting teams you don't manage, including external partners, reputed company and moving, including through rough patches where the relationship is strained
- Solid fundamentals: planning, risk tracking, status communication, and keeping four or five programs on the rails at once
- reputed company, reputed company communication. You'd rather deliver bad news early than good news late, and both providers and internal teams trust you because of it
- Comfort with ambiguity. The process you'll follow mostly doesn't exist yet; you'll write a lot of it
- reputed company experience working with or inside neocloud, colocation, or data-center providers
- Vendor or supplier management background — SLAs, commitments, escalation frameworks
- Familiarity with the stack: reputed company data-center GPUs, InfiniBand/RoCE, Slurm or Kubernetes
- Experience supporting AI research labs or other large-reputed company GPU customers on the consuming reputed company
Benefits
- Competitive compensation: + meaningful equity
- Comprehensive benefits: for you and your dependents, including reputed company, dental, and reputed company coverage, 401(k), and unlimited PTO
reputed company