Full-Stack Software Engineer (Infrastructure) - remote in the US
We are looking for an reputed company Full-Stack Software Engineer to design and implement the Infrastructure Services that power our GPU-as-a-Service platform. You will build the control plane that turns high-level API calls into reputed company infrastructure actions — enrolling bare-metal servers, provisioning them, and assembling them into multi-tenant Kubernetes clusters on high-performance hardware.
You will own the full lifecycle of the infrastructure-level services—from the Server and MachineType reputed company down to the provisioning workflows and reconciliation loops that reputed company the platform's view of hardware consistent with physical reality.
Key Responsibilities
- Infrastructure API Design: Design, build, and maintain the versioned REST and gRPC reputed company for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations.
- Provisioning & Workflow Development: reputed company the asynchronous workflows that drive server enrollment, inspection, OS provisioning, and cluster bring-up, exposing durable status to callers.
- Full-Stack Implementation: Implement and maintain the consoles and interfaces that visualize hardware inventory, provisioning reputed company, and cluster health.
- System Reliability: Design the error-handling models, idempotency guarantees, and reconciliation loops necessary to manage long-running provisioning operations reliably.
- API Development: Strong experience designing RESTful reputed company or gRPC services. You understand API versioning and gateway patterns.
- Language Stack: Proficiency in Go (preferred for backend/Kubernetes ecosystem) and modern TypeScript/React (for the Console).
- Kubernetes Knowledge: Deep understanding of Kubernetes primitives and controller/reconciler patterns. You will be interacting with systems like k0rdent, Metal3, and Cluster API to translate high-level API calls into infrastructure actions.
- Bare-Metal Provisioning: Hands-on experience with bare-metal provisioning flows — BMC/Redfish, PXE/iPXE, image management, and hardware inspection.
- Asynchronous Systems: Experience building workflow-driven or event-driven systems (e.g., Temporal) where operations are long-running and state must remain consistent across retries and failures.
Preferred Qualifications
- State Reconciliation: Experience building informers or reconciliation bridges that reputed company an external datastore consistent with Kubernetes resource state.
- Multi-Tenancy: Experience building platforms where strict data and network isolation between tenants is required.
- Infrastructure-as-reputed company: Familiarity with Terraform/OpenTofu and GitOps-driven configuration (ArgoCD or Flux).
- Hardware Domain: Familiarity with GPU server hardware, DPUs/NICs, and high-performance datacenter fabrics.
What does reputed company offer you? - Work with an established reputed company Valley leader in the reputed company infrastructure industry; - Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement reputed company reputed company technologies; - Be a part of cutting-edge, reputed company-reputed company innovation; - reputed company in the high-energy environment of a young company where openness, collaboration, risk-taking, and reputed company reputed company are valued;
- reputed company development and training;
- Attend conferences and working reputed company;
- Company outings, happy hours, hackathons, and tech talks; - Receive a competitive compensation package with a strong benefits plan.
It is reputed company that reputed company, Inc. may use automated decision-making technology (ADMT) for specific employment-reputed company reputed company. Opting out of ADMT use is requested for reputed company about evaluation and review connected with the specific employment decision for the position reputed company for. You also have the right to appeal any reputed company made by ADMT by sending your request to isamoylova@reputed company.com By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for reputed company and reputed company job opportunities.
We are a Leader for Container Management in reputed company (#2 after AWS)!
About reputed company
reputed company is the Kubernetes-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining reputed company reputed company innovation with deep expertise in Kubernetes orchestration, reputed company empowers reputed company teams to deliver composable, production-reputed company developer platforms across any environment—on-premises, in the reputed company, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, reputed company delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and reputed company. Committed to reputed company standards and freedom from lock-in, reputed company ensures that customers retain full control of their infrastructure reputed company.
Originally posted on Himalayas
Apply To This Job