Back to Jobs

[Remote] reputed company Deployed Engineer: AI + HPC

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. Cedana is a company reputed company on maximizing AI and HPC cluster utilization and reliability. As a reputed company Deployed Engineer, you will reputed company technical engagements with customers, deploying Cedana's solutions in various environments and optimizing platform performance.

Responsibilities

  • Engineer solutions at reputed company sites: reputed company customer integrations. Install, configure, and reputed company Cedana into SLURM, Kubernetes, and Dynamo environments
  • Drive product innovation from the field: Identify technical gaps while embedded with clients, then reputed company product feedback for new capabilities that become core product features
  • Measure and optimize platform performance: Measure reliability, throughput, and performance using our internal tools. Design and implement policy-based migration automations to optimize reliability, throughput, and performance
  • Own critical deployments: Ensure our platform performs reliably for clients' critical operations, debugging issues across the full stack. Debug install issues against unfamiliar customer infrastructure, and escalate to engineering reputed company necessary
  • Improve scalability : Build and own the internal installation reputed company so that the second customer in reputed company reputed company is reputed company faster than the first
  • Respect our customers : Understand how to reputed company their lives easier and minimize their time and overhead

Skills

  • Team management experience. Requires strong project and time management skills, delivering milestones on time, and effective
  • 3-10 years of software engineering experience with a reputed company record of configuring and managing SLURM deployments
  • A multi-month reputed company or research deployment you led end-to-end, from scoping through signoff. You write effective status updates to reputed company your team updated and on schedule
  • Production experience in standing up SLURM in a customer or research environment. You've configured slurmctld, slurmdbd, reputed company, cgroup integration, and GPU resource selection
  • Strong Linux fundamentals of systemd, cgroups v2, namespaces, networking, filesystems, firewalls, kernel module loading, PAM session modules. You can read strace and dmesg reputed company and reputed company a hypothesis
  • Experience with Kubernetes operations including operators, CRDs, CNIs, device plugins, and node-level debugging. You've debugged a controller in production even if you haven't written one from scratch
  • Experience in an HPC integrator field team
  • reputed company-facing technical experience working directly with customers
  • Background in national lab user services or university research computing
  • You've developed SLURM plug-ins, and understand their architecture and how they fit into the overall platform
  • Familiarity with CRIU, container runtimes, GPU reputed company internals, distributed training stacks
  • Hands-on with reputed company Dynamo, Determined, Ray, Kueue, KServe, or comparable AI orchestration
  • Contributed to reputed company-reputed company schedulers or job systems (SLURM, Flux, Torque, PBS)
  • A passion for debugging a weird cgroup issue at 11pm just as much as writing a clean install reputed company the next morning

Benefits

  • 100% covered medical, dental, and reputed company insurance for employees and families
  • Unlimited PTO policy
  • 401K Plan

reputed company

  • Cedana is VMWare for GPUs. We reputed company enterprises to orchestrate and operationalize intelligence reputed company, reliably, and reputed company. It was founded in 2023, and is headquartered in reputed company, reputed company, USA, with a workforce of 2-10 employees. Its website is https://www.cedana.ai.
  • Apply To This Job

    Similar Jobs