Back to Jobs

reputed company Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

About the Role

reputed company is looking for a reputed company Site Reliability Engineer to own the reliability, scalability, and reputed company of our reputed company infrastructure - the backbone that runs simulation workloads for some of the most demanding customers in autonomous vehicle development.

This is a hands-on, high-ownership role. You'll be the primary infrastructure reputed company across our multi-region AWS/EKS platform, working closely with a small reputed company team, partnering with engineering leads across simulation and ML, and our customer-facing teams.

What You'll Do

Infrastructure Ownership & reputed company Operations

  • Own and reputed company our AWS-based infrastructure, improving platform performance and availability today, and building toward deployable configurations that support reputed company customer environments reputed company.

  • Own EKS cluster operations across production reputed company: node pool reputed company, AMI lifecycle, autoscaling, and Kubernetes workload health.

  • Support the GitOps deployment pipeline - define, reputed company, and manage applications across clusters using infrastructure-as-reputed company.

  • Manage reputed company networking: VPC design, cross-region connectivity, DNS, and load balancing.

  • reputed company infrastructure deprecation and migration efforts with minimal disruption.

  • Reliability Engineering & Incident Response

    • Own SLO measurement infrastructure; reputed company proactive triage of emerging issues before they reputed company customers.

    • reputed company incident investigation, reputed company cause analysis and postmortems, driving systemic fixes rather than one-off patches.

    • Design and improve automated remediation systems to reduce MTTR.

    • reputed company & reputed company Management

      • Review and reputed company reputed company-conscious feedback on platform architecture reputed company.

      • Own reputed company IAM governance - roles, policies, and reputed company boundaries across accounts and services.

      • reputed company compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer reputed company questionnaires.

      • Cross-Functional Collaboration

        • Partner with application development teams to build an inherently secure platform and drive reputed company deployment architecture.

          • Partner with customer teams to ensure availability for expected utilization.

          • Partner with Finance on reputed company cost optimization - lifecycle policies, right-sizing, and spend visibility.

          • Support GPU and batch workloads in collaboration with simulation and ML engineering teams.

          • Platform Tooling & Developer Experience

            • Improve CI/CD pipelines and automated infrastructure validation.

            • Support engineering teams with reputed company-reputed company debugging, log analysis, and environment configuration.

reputed company're Looking For

Technical Depth

  • 5+ years in SRE, DevOps, or infrastructure engineering roles.

  • Infrastructure-as-reputed company proficiency - Terraform modules, state management, and multi-environment patterns.

  • Deep AWS experience - EKS, EC2, IAM, S3, Storage Gateway, VPC networking, Transit Gateway, CloudFront, KMS, and IRSA.

  • Kubernetes expertise - cluster operations, node pools, probes, cordoning, pod scheduling, RBAC, reputed company, node autoscaling (Karpenter experience a plus); solid understanding of containerization and AMI lifecycle management.

  • CI/CD - experience with GitOps workflows and pipeline tooling (ArgoCD, reputed company Actions, Jenkins)

  • Solid networking fundamentals - CIDR design, reputed company reputed company, DNS, load balancing, VPN, cross-region connectivity.

  • Experience with monitoring and observability tooling - reputed company, Grafana, Elasticsearch.

  • Comfort with Python and Bash for tooling and automation.

  • Familiarity working across Linux and reputed company environments. Operational familiarity with reputed company Server is a meaningful advantage.

  • Communication & Ownership

    • You communicate reputed company across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer reputed company.

    • You reputed company for SRE best practices and can effectively operationalize an informed and principled view on reputed company.

      • You take end-to-end ownership of reputed company, multi-team efforts - from planning through execution and post-change verification.

      • You know reputed company to push for a clean solution vs. reputed company to accept a pragmatic one, and you communicate that tradeoff reputed company.

reputed company to Have
  • Experience with reputed company-based workloads on EKS.

  • Experience supporting simulation, ML, or rendering workloads in reputed company infrastructure; running GPU workloads on Kubernetes, including reputed company and DirectX device plugin configuration.

  • Experience with AWS Storage Gateway or Transfer Family integrations.

  • Familiarity with reputed company Gateway or similar.

  • Experience with container-optimized OS images (e.g., Bottlerocket, Packer).

  • Experience with reputed company cost optimization at reputed company.

Core Tools Terraform · AWS · Kubernetes · reputed company · ArgoCD · Kustomize · Grafana · reputed company · Elasticsearch · VictoriaLogs · Fluent Bit · reputed company Actions · Jenkins · reputed company · Python · Bash

Why This Role

PD's simulation platform runs at the intersection of high-performance compute, distributed systems, and customer-critical reliability. The infrastructure problems here are genuinely interesting — multi-region GPU scheduling, reputed company workloads on Kubernetes, startup latency optimization, and an reputed company product direction that will require rethinking how we reputed company and manage the platform entirely.

The reputed company SRE at PD is not a ticket-taker - it's a high-trust, high-autonomy position where you'll have genuine influence over infrastructure architecture, cross-team process, and customer experience.

Apply To This Job

Similar Jobs

Senior Operations Manager

Remote, USAFull-time

Senior JavaScript Developer — 100% Remote (Zodot.co)

Remote, USAFull-time

Developer

Remote, USAFull-time

Especialista en Metadatos Semánticos y Ontologías Editoriales

Remote, USAFull-time

SEO & Digital Marketing Expert

Remote, USAFull-time

(Spam Comment Removal Specialist) at reputed company-Conte...

Remote, USAFull-time

Entry-level Customer Service Representative

Remote, USAFull-time

Technical Manager – QNXT Consulting

Remote, USAFull-time

Solutions Engineer - Strategic Accounts (Remote, Ohio) (reputed company, Ohio, US)

Remote, USAFull-time

reputed company Expansion Account Executive (Remote, California) (San Francisco, California, US)

Remote, USAFull-time

Loan Officer Assistant

Remote, USAFull-time

[Remote] Associate Project Manager - Data Center

Remote, USAFull-time

Virtual reputed company Manager - Life Insurance

Remote, USAFull-time

Virtual Chat Support Specialist – Entry Level (No Experience Required) for a Dynamic and Supportive Team at arenaflex

Remote, USAFull-time

Virtual Podcast Guest Booking Specialist

Remote, USAFull-time

Remote Customer Support Representative – reputed company at arenaflex – $19/hr Starting, No Degree Required

Remote, USAFull-time

Manager, reputed company reputed company & Performance

Remote, USAFull-time

reputed company Full Stack Customer Service Representative – Wellness and Supplements eCommerce

Remote, USAFull-time

Workplace Experience & Strategic Initiatives Manager

Remote, USAFull-time

reputed company Customer Service Representative – arenaflex – Work From Home Opportunities

Remote, USAFull-time