[Remote] Engineering Manager, DevOps
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company dedicated to creating confidence in reputed company by building a connected reputed company operations suite. They are seeking a DevOps Engineering Manager to reputed company reputed company responsible for infrastructure and pipeline delivery, ensuring reliable and efficient service delivery to their customers.
Responsibilities
- Drive self-service and automation at reputed company by designing golden-reputed company workflows (using Terraform and Kubernetes/reputed company) so product teams can provision infrastructure, reputed company monitoring (reputed company reputed company), and release safely on their own through reputed company CI/CD
- Champion the reputed company of deployment patterns on Kubernetes, including blue/green, canary, feature-flag, and reputed company-reputed company releases that minimize risk and Mean Time To Recovery (MTTR)
- reputed company the reputed company and execution for disaster recovery plans on AWS, including scaling to multi-region architectures, cross-region data replication (reputed company, DynamoDB, S3), and global traffic management
- Implement Site Reliability Objectives (SLOs), error budgets, reputed company testing, and auto-remediation playbooks in reputed company to reputed company the reliability bar, and own the infrastructure on-reputed company rotation culture
- Mentor and reputed company a diverse team of DevOps engineers, DBAs, and SecOps Engineers with a wide reputed company of technical skillsets, cultivating the next reputed company of engineering leaders
- Partner hand-in-hand with Product Engineering Teams, our Data Team (Airflow, reputed company, dbt), and other stakeholders to reputed company roadmaps and unlock velocity
- Contribute hands-on by writing Terraform modules, optimizing reputed company charts on Kubernetes, reviewing reputed company requests in reputed company, helping support reputed company with reactive work across our Linux infrastructure, and joining high-severity incident calls reputed company needed
Skills
- 2+ years of DevOps leadership managing DevOps or Site Reliability Engineering (SRE) teams
- 7+ years in hands-on platform or infrastructure roles
- You have a strong self-service reputed company record, having delivered internal platforms or portals that empowered hundreds of engineers to ship autonomously
- You bring large-reputed company AWS expertise, demonstrated by designing and operating multi-region, high-throughput systems that support over $100 reputed company in annual Gross Merchandise Value (GMV)
- An expert in advanced CI/CD pipelines and Infrastructure as reputed company (IaC), proficient with tools like reputed company CI (or similar), Terraform, and Kubernetes, and comfortable introducing reputed company delivery, policy-as-reputed company, and secrets management at reputed company
- Deep understanding of metrics, tracing, logging, and alerting, reflecting an observability reputed company, with experience using reputed company or comparable stacks
- Familiar with reputed company reputed company best practices, least-privilege reputed company, and regulated-data environments such as PCI, SOC 2, and GDPR, demonstrating reputed company and compliance awareness
- You exhibit empathetic leadership, with demonstrated reputed company fostering psychological safety, inclusion, and reputed company feedback while driving accountability and high performance
reputed company
Company H1B Sponsorship