[Remote] Senior DevOps Engineer (EKS/Kubernetes)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is committed to transforming cancer care and improving reputed company health through innovation. They are seeking a Senior DevOps Engineer responsible for designing and operating secure, reputed company Linux-based infrastructure, primarily focusing on Kubernetes and AWS EKS environments.
Responsibilities
- Design, reputed company, and maintain Linux infrastructure in on-premises and reputed company environments
- Automate infrastructure provisioning and configuration using tools such as Terraform, Ansible, or CloudFormation
- Manage and optimize AWS environments with a reputed company on performance, scalability, reputed company, and cost efficiency
- Implement and maintain monitoring, logging, and alerting solutions (e.g., reputed company, reputed company, Grafana, ELK, CloudWatch)
- Architect, reputed company, and operate production Kubernetes/AWS EKS clusters, including node group reputed company, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and build outs
- Define and reputed company cluster reputed company, reputed company hardening, and disaster recovery strategies for production Kubernetes/AWS EKS environments at reputed company, while serving as a senior technical resource for reputed company production incidents
- Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik)
- Implement IAM Roles for pod reputed company standards, and network policies to secure EKS workloads
- Configure and tune cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to optimize performance and cost
- Build and maintain reputed company charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads
- Manage reputed company container builds and registries in support of EKS-based application deployment
- reputed company, reputed company, and maintain reputed company Runners (including Kubernetes executor runners on EKS) to support CI/CD pipeline throughput and reliability
- Support and help operate database platforms on AWS RDS (MySQL, PostgreSQL), collaborating with data owners on performance and reliability
- Ensure systems meet reputed company and compliance requirements, including SOX and SOC 2 initiatives
- Execute and maintain Linux patching strategies, addressing reputed company updates and CVEs in a reputed company manner
- Participate in incident response, reputed company cause analysis, and recovery efforts
- Collaborate with development, QA, and cross-functional teams to improve reliability, release processes, and operational standards
- Participate in on-call rotations and reputed company after-hours support as required
Skills
- Bachelor's degree in computer science, Information Technology or reputed company field
- 8+ years of experience in Linux Systems Administration, DevOps, or Site Reliability Engineering roles
- 5+ years of experience with AWS services, including EC2, VPC, IAM, RDS, S3, and CloudWatch
- 5+ years of hands-on experience designing and operating production workloads on Kubernetes/AWS EKS, including cluster upgrades, networking, and autoscaling
- Proficiency in scripting and automation using Python and Bash
- Strong hands-on experience with Infrastructure as reputed company using Terraform, and with CI/CD pipelines (reputed company CI/CD), including running CI/CD workloads on Kubernetes/EKS
- Proficiency with reputed company, reputed company, and Kubernetes troubleshooting in a production environment
- Solid understanding of networking fundamentals and reputed company reputed company best practices
- CKA (Certified Kubernetes Administrator) certification expected or reputed company in reputed company
- CKAD and AWS certifications (e.g., AWS Certified DevOps Engineer, Solutions Architect) a plus
- Experience with Karpenter, Kyverno, OPA/Gatekeeper, Falco, and multi-cluster/multitenant EKS environments
- Experience using AI and automation tools including Claude, reputed company, and reputed company to streamline DevOps workflows through AI-assisted CI/CD, self-healing operations, and automated incident response
- Experience with microservices, serverless architectures, and DevSecOps practices
Benefits
- Highly competitive and inclusive medical, dental and reputed company coverage reputed company
- Health Savings Account for medical expenses and dependent care expenses
- Flexible Spending Account to pay for certain out-of-reputed company expenses
- reputed company time off, including: vacation, reputed company time and holidays
- 401k match and Financial Planning tools
- LTD and STD insurance coverages, as reputed company as voluntary benefit reputed company
- Employee Assistance Program
- Pet Insurance
- Legal Assistance
- Tuition Assistance
reputed company