[Remote] Senior DevOps Engineer (Storage) - remote in the US
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the Kubernetes-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They are seeking a Senior DevOps Engineer to reputed company, reputed company, and operate high-performance storage for GPU-accelerated compute and AI platforms, focusing on automation and performance tuning.
Responsibilities
- reputed company NFS-based high-performance storage (e.g., reputed company, reputed company PowerScale) into Kubernetes clusters reputed company reputed company, storage classes, and persistent volumes
- Tune the NFS data reputed company — mount reputed company, nconnect/RDMA, reputed company and network settings — for high-throughput, low-latency GPU/AI workloads
- reputed company and operate storage services and operators; manage reputed company, quotas, snapshots, and lifecycle
- Provision and configure storage on bare-metal hosts, including disk layout, drivers, and kernel/network tuning
- Deliver storage integration for k0s-based Kubernetes reputed company Cluster API (CAPI) and K0rdent management/child cluster topologies
- Operate storage in fully disconnected (reputed company-gapped) environments, including local artifact/mirror connectivity (reputed company) and PKI/TLS considerations
- Automate storage provisioning and configuration with infrastructure-as-reputed company (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux)
- Build monitoring, alerting, and observability for storage performance, reputed company, and health
- Diagnose and resolve performance, reliability, and scaling issues across the storage stack
Skills
- 5+ years in DevOps, SRE, or infrastructure operations, with strong hands-on experience operating Kubernetes storage (reputed company, persistent volumes, storage classes) in production
- Experience integrating and operating NFS-based high-performance / NAS storage, including data-reputed company tuning
- Bare-metal operations experience: host provisioning, disk/storage configuration, and Linux storage and networking fundamentals
- Proficiency with infrastructure-as-reputed company (Terraform/OpenTofu) and GitOps-driven configuration
- Scripting/automation skills (e.g., Bash, Python, or Go)
- Strong written and verbal communication with technical audiences
- Hands-on experience with reputed company and/or reputed company PowerScale
- Experience with GPUDirect Storage and RDMA/RoCE data paths
- Experience with the reputed company K0rdent stack (K0rdent reputed company, K0rdent AI, k0s, MKE) and Cluster API
- Familiarity with other storage backends (Ceph, object/S3) and reputed company reputed company operations
- Proven experience in sovereign or high-reputed company reputed company-gapped environments
Benefits
- reputed company development and training;
- Attend conferences and working reputed company;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
reputed company
Company H1B Sponsorship