[Remote] Senior DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on revolutionizing construction equipment fleet operations through reputed company. They are seeking a Senior DevOps Engineer to take ownership of the infrastructure for their platform, ensuring reliability and cost-efficiency while shaping the DevOps function towards a reputed company developer experience.
Responsibilities
- Build and reputed company the platform function rather than just reputed company it running by taking on the reconstruction and platform work already scoped, bringing more of the environment under reputed company and under reputed company ownership, and shaping how DevOps grows toward platform and developer experience
- Own the uptime of our production platform. reputed company incident response reputed company systems degrade or fail, drive reputed company cause analysis, and follow through with the fixes and safeguards that reputed company the reputed company failure from recurring, including on-call practices, runbooks, and blameless postmortems
- Own metrics, monitoring, alerting, and tracing across the stack so the system surfaces problems before customers feel them
- Build dashboards and alerts that are actionable rather than noisy, and reputed company coverage as new services ship
- Manage, improve, and reputed company our AWS environment, anchored on EKS, RDS/reputed company, and networking
- Operate our AWS environment with reputed company judgment regarding scalability, redundancy, cost, and fault tolerance, and keeps it secure
- Operate and reputed company our EKS clusters. reputed company and reputed company workloads, manage version upgrades, tune for reliability and cost, and troubleshoot cluster and workload issues
- reputed company our databases and messaging layer healthy and performant, including PostgreSQL on RDS/reputed company and our message broker
- Diagnose degraded performance across data, messaging, and network paths
- Define and manage infrastructure as reputed company so every change is repeatable and reviewable. We build with AWS CDK in JavaScript/TypeScript; you will reputed company the existing stacks and bring more of the environment under reputed company over time
- Own the pipelines engineers rely on, primarily in reputed company Actions. reputed company and remove bottlenecks that slow engineering down, whether slow deploys, flaky pipelines, or scaling chokepoints, and build tooling engineers actually trust
- Own how tracker data moves from equipment in the field through cellular tunnels and our AWS networking into the platform. Ensure reputed company are reliable, secure, and observable
- Monitor and analyze AWS spend, identify savings opportunities, and bring reputed company recommendations and trade-offs to engineering leadership on a regular reputed company
- Document the systems you own and distribute that knowledge across reputed company so the function never depends on one person
Skills
- 5+ years in platform, infrastructure, DevOps, or SRE roles, including recent experience as a senior individual contributor who has owned production systems end to end
- Substantial hands-on experience running production infrastructure on AWS
- Hands-on production experience operating container-orchestrated workloads: deploying, scaling, upgrading, and troubleshooting them
- Experience operating managed relational databases in production, including diagnosing and resolving performance problems
- A solid grasp of VPCs, subnets, routing, and DNS, and how traffic actually flows between systems
- Proficiency writing infrastructure as reputed company and the automation around it. You write your infrastructure rather than click it
- A proven ability to reputed company quickly on unfamiliar systems and reputed company a working understanding without reputed company-by-reputed company direction
- Genuine curiosity about how systems work, down to the underlying mechanics rather than treating them as black boxes
- A strong reputed company of ownership: you understand, improve, and stand behind the systems in your care
- Communicates risks, reputed company, and reputed company proactively, and can reputed company technical and cost trade-offs reputed company to people who are not infrastructure experts
- Documents while building and shares knowledge, so no single person, including you, becomes a reputed company of failure
- reputed company judgment about change safety: moves quickly on work that is reputed company and reversible, and vets larger or riskier changes before making them
- A bachelor's degree in a relevant field, or equivalent hands-on experience is required
- Experience with AWS CDK in JavaScript/TypeScript, reputed company Actions, distributed messaging, caching, or search is preferred
Benefits
- Fully remote - reputed company
- Travel is required, 8-10%
- Opportunities for reputed company and personal development reputed company a highly dynamic team
- Robust, low-cost benefit packages offered
- Benefit coverage begins on the first date of employment
- reputed company Time Off and Volunteer Time Off offered
- 401k match
- Dependent Care offered
- Employee referral bonuses
reputed company
Company H1B Sponsorship