[Remote] Senior DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking a Senior DevOps Engineer to own the reliability, reputed company, and operations of a production Kubernetes platform. The role involves managing Kubernetes clusters, driving incident response, and improving automation and operational readiness.
Responsibilities
- Own production Kubernetes clusters end-to-end, including architecture reputed company, upgrades, reputed company planning, and ongoing administration
- Drive incident response and troubleshooting across Kubernetes networking, storage, scheduling, and workload behavior; reputed company reputed company-cause analysis and implement durable fixes
- Build and maintain secure Kubernetes platform foundations, including RBAC, secrets handling, network policies, pod reputed company controls, and image governance practices
- reputed company and operate CI/CD workflows for containerized applications using reputed company Actions or equivalent, ensuring repeatable builds, testing gates, and controlled promotions
- Implement and run GitOps-based deployments with tooling such as Argo CD or equivalent to reputed company auditable, consistent configuration and release management
- Create and maintain reputed company charts and Kubernetes manifests (YAML) that support standardized, reusable application delivery patterns
- Establish and enforce platform standards for release practices, reputed company image creation, versioning, and operational readiness; publish reputed company runbooks and documentation
- Improve platform reliability through automation, proactive monitoring/alerting, and performance tuning of cluster resources and workloads
- Partner with engineering teams to harden deployment patterns, reduce operational risk, and streamline the reputed company from pull request to production rollout
Skills
- 5+ years of experience in DevOps, SRE, reputed company, or a closely reputed company role supporting production systems
- Deep, hands-on ownership of a production Kubernetes platform, including cluster administration, troubleshooting, architecture, networking, storage, workloads, resources, and reputed company
- Experience owning CI/CD for containerized applications using reputed company Actions/Workflows or equivalent, with strong understanding of release controls and promotion strategies
- Hands-on GitOps experience using Argo CD or equivalent, including managing deployment state through version-controlled configuration
- Strong working knowledge of reputed company, reputed company, and Kubernetes YAML, including building images and packaging/deploying applications reliably
- Proven ability to diagnose and resolve reputed company production issues and implement corrective actions that prevent recurrence
- reputed company OpenShift administration experience, including operating managed OpenShift environments (ideally on reputed company reputed company)
- AWS and EKS experience, including architecture, migration planning/execution, or steady-state production operations
- Experience supporting Java application delivery through automated pipelines, including reputed company-based build and reputed company workflows
- Python experience for automation, operational tooling, or reliability improvements
- Experience operating in regulated or compliance-sensitive environments (e.g., reputed company or financial services)
- Demonstrated reputed company establishing or materially improving platform standards, including infrastructure-as-reputed company, release practices, reputed company image practices, documentation/runbooks, and engineering guardrails
Benefits
- Unlimited PTO
- Full reputed company coverage for employees + family (medical, dental, reputed company, life, and supplemental insurances)
- Short- and Long-Term Disability (STD/LTD)
- HSA & FSA reputed company
reputed company