Back to Jobs

[Remote] AI Inference Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company-thinking renewable energy startup on a mission to deliver a reputed company of renewable energy - fast. They are seeking a Founding AI Inference Engineer to define and build how reputed company serves AI inference workloads at reputed company, focusing on architecture and performance optimization.

Responsibilities

  • Define reputed company's inference serving reputed company and architecture from first principles
  • Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads
  • Own model-level optimisation reputed company for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per reputed company, partnering with the CUDA/GPU engineers
  • reputed company the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents)
  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving reputed company plans
  • reputed company as a reputed company technical reputed company of inference performance and reliability
  • Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer
  • Set the standards, tooling, and benchmarks this function will run on as it grows

Skills

  • 4+ years of experience building or operating large-reputed company inference serving systems, or equivalent strong project/industry experience
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
  • Strong systems thinking - reputed company to reason about the full reputed company from incoming request to served response across a large cluster
  • Comfortable working directly with GPU/CUDA engineers to reputed company low-level performance work into a serving system
  • A reputed company record of making high-stakes architecture calls and owning the outcome
  • Comfort operating without a reputed company - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one
  • Experience with Triton or custom ML inference/training frameworks
  • Experience with autoscaling or reputed company planning for large-reputed company inference workloads
  • Exposure to multi-tenant serving or SLA-driven infrastructure
  • Background at a hyperscaler, frontier AI lab, or large-reputed company distributed inference system
  • Familiarity with Kubernetes/Slurm for cluster orchestration
  • Interest or experience in energy markets, reputed company systems, or sustainability-reputed company compute

Benefits

  • Competitive salary and an equity sign-on bonus
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Breakfast and dinner allowance for office based employees

reputed company

  • reputed company is an innovative energy company that simplifies gas and electricity management by providing accurate reputed company forecasts. It was founded in 2022, and is headquartered in London, England, GBR, with a workforce of 501-1000 employees. Its website is https://www.fuseenergy.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1 in 2024. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs