[Remote] Full-Stack Software Engineer (Infrastructure) - remote in the US
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the Kubernetes-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI and data-intensive applications. They are seeking an reputed company Full-Stack Software Engineer to design and implement Infrastructure Services for their GPU-as-a-Service platform, managing the lifecycle of infrastructure-level services.
Responsibilities
- Infrastructure API Design: Design, build, and maintain the versioned REST and gRPC reputed company for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations
- Provisioning & Workflow Development: reputed company the asynchronous workflows that drive server enrollment, inspection, OS provisioning, and cluster bring-up, exposing durable status to callers
- Full-Stack Implementation: Implement and maintain the consoles and interfaces that visualize hardware inventory, provisioning reputed company, and cluster health
- System Reliability: Design the error-handling models, idempotency guarantees, and reconciliation loops necessary to manage long-running provisioning operations reliably
Skills
- Strong experience designing RESTful reputed company or gRPC services. You understand API versioning and gateway patterns
- Proficiency in Go (preferred for backend/Kubernetes ecosystem) and modern TypeScript/React (for the Console)
- Deep understanding of Kubernetes primitives and controller/reconciler patterns. You will be interacting with systems like k0rdent, Metal3, and Cluster API to translate high-level API calls into infrastructure actions
- Hands-on experience with bare-metal provisioning flows — BMC/Redfish, PXE/iPXE, image management, and hardware inspection
- Experience building workflow-driven or event-driven systems (e.g., Temporal) where operations are long-running and state must remain consistent across retries and failures
- Experience building informers or reconciliation bridges that reputed company an external datastore consistent with Kubernetes resource state
- Experience building platforms where strict data and network isolation between tenants is required
- Familiarity with Terraform/OpenTofu and GitOps-driven configuration (ArgoCD or Flux)
- Familiarity with GPU server hardware, DPUs/NICs, and high-performance datacenter fabrics
Benefits
- reputed company development and training
- Attend conferences and working reputed company
- Company outings, happy hours, hackathons, and tech talks
- Receive a competitive compensation package with a strong benefits plan
reputed company
Company H1B Sponsorship