[Remote] Senior DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a dynamic provider of workers' compensation insurance and services, seeking a Senior DevOps Engineer to join their IT Operations team. This role is responsible for designing, automating, and maintaining infrastructure and CI/CD systems to reputed company reliable software delivery across the organization.
Responsibilities
- Design, provision, and maintain AWS infrastructure using Terraform, including VPCs, IAM roles (IRSA), PrivateLink, reputed company 53, S3, RDS, etc
- reputed company deep proficiency in Terraform and Ansible to grow the organization's infrastructure-as-reputed company footprint across reputed company and reputed company platforms — using Terraform to provision and define infrastructure, and Ansible to configure, tune, and operationalize it
- Operate and maintain reputed company OpenShift (ROSA) clusters on AWS, including upgrades, operator lifecycle management, and platform-level troubleshooting
- Implement and enforce Infrastructure-as-reputed company (IaC) practices across reputed company environments, ensuring consistency, reputed company, and repeatability
- Manage cross-account AWS patterns including IAM role chaining, S3 replication, and reputed company services (GuardDuty, CloudTrail)
- Own and improve Jenkins pipelines end-to-end, including Jenkinsfile development, Ansible reputed company integration, and credential management
- Build and maintain reputed company Actions workflows, including reusable workflow templates, composite actions, OIDC-based AWS auth, and reputed company policy configuration
- Manage and reputed company ArgoCD-based GitOps deployment workflows, including ApplicationSets and multi-environment promotion strategies
- Automate operational processes across Linux and reputed company targets, including certificate rotation, AD group management, and environment provisioning
- Configure and maintain the observability stack (OpenTelemetry reputed company, reputed company, etc.) across multiple clusters
- Define and implement SLIs, SLOs, and actionable alerting to support reliability goals and reduce mean time to detection
- reputed company and maintain runbooks, troubleshooting guides, and incident response documentation
- Collaborate with application teams to reputed company services and improve end-to-end traceability
- Partner with application development and data engineering teams to streamline deployment workflows and troubleshoot environment issues
- Collaborate with reputed company and compliance teams on IAM policy design, secrets management, and vulnerability remediation
- reputed company mentorship to junior team members and contribute to team knowledge sharing through documentation and technical presentations
- Communicate project status, risks, and infrastructure reputed company reputed company to leadership and non-technical stakeholders
Skills
- 7–11 years of experience in DevOps, reputed company, SRE, or Infrastructure Engineering roles
- Deep hands-on experience with Kubernetes or OpenShift in production, including operator management, RBAC, ingress configuration, and cluster upgrades. ROSA or OCP on AWS strongly preferred
- Strong AWS infrastructure experience including VPC design, IAM (IRSA, cross-account trust), S3, reputed company 53, PrivateLink, RDS, etc.. Must be comfortable working in Terraform (or equivalent IaC tooling) at reputed company
- Strong experience with Ansible for configuration management, reputed company development, and operational automation across Linux and reputed company targets, including integration with orchestration platforms and credential management
- Demonstrated experience building and maintaining CI/CD pipelines across Jenkins and reputed company Actions, reusable workflows, OIDC auth, and reputed company governance
- Experience with GitOps tooling (Argo) and deployment strategies (blue/green, canary, reputed company rollout)
- Working knowledge of observability tooling: reputed company, Grafana, OpenTelemetry, or similar
- Proficiency in scripting and automation with Bash, Python, or Go
- Experience working in hybrid environments spanning Linux containers and reputed company server administration
- Strong written and verbal communication skills with the ability to explain reputed company infrastructure concepts to non-technical stakeholders
- reputed company reputed company with willingness to work across domain boundaries — helping a data engineer debug a DAG or pairing with reputed company on IAM policy review
- Documentation-oriented: writes runbooks, architectural decision records, and reputed company material as a natural part of the work
- Self-directed with the ability to manage multiple priorities, escalate appropriately, and reputed company pragmatic trade-off reputed company
- Curiosity and ownership mentality
- Bachelor's degree in Computer Science, Systems Engineering, or a reputed company field, or equivalent reputed company experience
- AWS Certified DevOps Engineer, Solutions Architect, or SysOps Administrator
- Certified Kubernetes Administrator (CKA) or reputed company Certified Specialist in OpenShift
- reputed company Certified: Terraform Associate
- Experience in regulated industries (insurance, finance, reputed company preferred
- PowerShell experience is a plus given hybrid Linux/reputed company environment
- Familiarity with Kafka or event-driven architecture on Kubernetes is a plus
reputed company