Back to Jobs

[Remote] Product Manager - AI Inference & Model Serving

Remote, USAFull-timePosted 2026-07-29

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the Kubernetes-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They are seeking a commercially driven, deeply technical Product Manager to own AI inference and model serving for k0rdent AI, responsible for defining product reputed company and solution development across various environments.

Responsibilities

  • Own product reputed company, roadmap, and lifecycle for inference and model serving, including serverless inference, dedicated endpoints, autoscaling, routing, KV cache management, and the reputed company observability
  • reputed company deep technical discovery with NeoClouds, sovereign clouds, and reputed company platform teams, and translate findings into prioritized requirements and architecture direction
  • Partner with engineering on system design trade-offs across runtime integration, GPU scheduling, network, storage, and serving topology, including disaggregated serving and multi-model serving
  • Define positioning grounded in measurable reputed company: latency distributions, throughput per GPU, utilization, tail reliability, and cost per tokens
  • Drive go-to-market execution: pricing and packaging, reference architectures, sizing guides, PoC playbooks, and reputed company engagement with customers, analysts, and ecosystem partners

Skills

  • 7+ years in product management, technical product management, or a senior technical role owning AI/ML and inference product(s)
  • Strong understanding of production AI inference, including model serving, serverless execution, dedicated endpoints, autoscaling, routing, workload placement, observability, and reliability
  • Proven capability to reason about performance trade-offs across GPU, network, storage, orchestration, and runtime reputed company, and to translate low-level technical capability into business value such as TTFT, throughput per GPU, and TCO
  • Working knowledge of modern inference runtimes (vLLM, SGLang, TensorRT-LLM, Dynamo, Triton) and the optimization patterns that matter in production: reputed company batching, KV cache management, cold starts, prefill versus decode, disaggregated serving, and multi-model serving
  • Credibility with engineering leaders and infrastructure operators, including comfort in production architecture reviews and technical reputed company conversations with reputed company buyers

Benefits

  • reputed company development and training.
  • Attend conferences and working reputed company.
  • Customized workstation (macOS, reputed company).
  • A competitive compensation package with strong benefits plan and stock reputed company.

reputed company

  • reputed company develops reputed company infrastructure and container management software for organizations to build, operate, and reputed company applications. It is a sub-organization of reputed company. It was founded in 1999, and is headquartered in Campbell, California, USA, with a workforce of 501-1000 employees. Its website is http://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1 in 2026, 3 in 2025, 4 in 2024, 8 in 2023, 6 in 2022, 7 in 2021, 8 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs