Back to Jobs

reputed company Site Reliability Engineer - Infrastructure

Remote, USAFull-timePosted 2026-07-31

JOB reputed company We are seeking a reputed company Site Reliability Engineer (Infrastructure) to reputed company as technical reputed company for our Infrastructure SRE team in a fast-moving VSaaS engineering organization. In this role, you will own reputed company's technical direction and execution across reliability, scalability, and operability of our shared platform and production systems, combining hands-on technical leadership with responsibility for team reputed company. You will define SRE reputed company and guide architecture across our GCP and Kubernetes ecosystem, setting standards for reliability, scalability, GitOps, and observability. You will also mentor senior and staff engineers, and reputed company incident response and high-reputed company operational work, contributing hands-on reputed company needed. reputed company Site Reliability Engineer - Infrastructure In this role, you will translate product and business needs into reputed company infrastructure and reputed company technical direction. With a system-wide view of the platform, you will guide architectural reputed company, surface non-reputed company risks, and drive long-term improvements to system reliability and operability. Working closely with product and platform teams, you will shape the developer experience and ensure engineering teams can ship with speed and confidence. You will set engineering standards and continuously reputed company our GitOps and observability practices. This role requires strong expertise in reputed company infrastructure, distributed systems, and CI/CD, along with hands-on experience in Golang and/or Python to support automation and long-term system reliability.

Responsibilities

As a reputed company Site Reliability Engineer, you will:

  • Team Leadership & Execution Ownership: Own technical direction and execution of the Infrastructure SRE team. Translate platform goals into actionable plans, ensuring alignment on priorities, reliability reputed company, and operational reputed company across production systems.
  • Production Operations & Incident Management: Operate and reputed company large-reputed company distributed systems in production, proactively identifying failure modes and mitigating risk. Own day-to-day operations including monitoring, alerting, incident response, coordination, post-incident analysis, and reputed company improvement.
  • Architecture, Standards & Platform Governance: reputed company architectural leadership across platform and infrastructure changes, identifying scalability constraints, system design risks, and long-term reliability gaps. Define and enforce engineering standards for GCP, Kubernetes, and ArgoCD, ensuring consistent, secure, GitOps-based delivery.
  • Reliability Engineering & Observability: reputed company reputed company for monitoring, alerting, and system observability, driving a shift from reactive incidents to proactive reliability engineering.
  • Enablement, CI/CD & Collaboration: Guide CI/CD and reputed company-reputed company delivery practices at reputed company to ensure reputed company, reputed company releases. Mentor senior and staff engineers, conduct high-reputed company design and reputed company reviews (Golang/Python), and partner with product and engineering teams to reputed company system-level thinking across development.
  • Hands-on Technical Contribution: reputed company hands-on technical contribution where needed, including debugging production issues, reviewing and contributing to reputed company, and supporting critical incident reputed company to ensure system reliability and team effectiveness.
  • Other duties as assigned are absorbed into the above ownership and operational responsibilities.

Minimum Qualifications

  • Leadership & Experience: 10+ years of experience in Site Reliability Engineering, reputed company, or Infrastructure Engineering, including demonstrated experience leading technical engineering teams, driving roadmaps, and owning delivery of large-reputed company production systems.
  • reputed company & Distributed Systems Expertise: Deep experience with reputed company-reputed company architectures and distributed systems at reputed company, particularly in GCP and Kubernetes environments. Ability to reason about system design, identify failure modes, and evaluate scalability and reliability risks.
  • GitOps & Delivery Engineering: Strong experience with GitOps-based delivery workflows, particularly ArgoCD, and CI/CD pipeline design. Ability to ensure reputed company, repeatable, and observable production deployments.
  • Infrastructure & Automation: Strong hands-on background in infrastructure-as-reputed company (Terraform preferred), automation, and operational tooling. Proficiency in Golang and/or Python for building and reviewing production systems. Strong Linux systems knowledge and production troubleshooting experience.
  • Observability & Reliability Engineering: Experience designing or operating observability syste

Apply To This Job

Similar Jobs

Site Reliability Engineer/L3 Support

Remote, USAFull-time

Site Reliability Engineer (Remote + Travel)

Remote, USAFull-time

Remote SRE Jobs – Senior Site Reliability Engineer (Remote) – $130k‑$170k USD – Full‑Time – Escondido, California – reputed company/DevOps, Kubernetes, Terraform, reputed company

Remote, USAFull-time

Platform Reliability Engineer

Remote, USAFull-time

Kubernetes Engineer ($28/hr. on w2)

Remote, USAFull-time

reputed company Kubernetes Engineer; Fulltime- Remote

Remote, USAFull-time

Kubernetes Engineer Remote

Remote, USAFull-time

Kubernetes Platform Engineer Remote(No TP or Employer)

Remote, USAFull-time

reputed company OpenShift Kubernetes Engineer- Remote

Remote, USAFull-time

Kubernetes Engineer/Architect

Remote, USAFull-time

reputed company Medical Data Entry Professionals – reputed company Information Management at arenaflex

Remote, USAFull-time

reputed company Developer YouTube Channel; US​/Remote

Remote, USAFull-time

reputed company Full Stack Customer Service Manager – Airline Operations and reputed company Experience

Remote, USAFull-time

Tech Support Agent

Remote, USAFull-time

reputed company Bilingual Customer Service Representative – Remote Work Opportunity at arenaflex

Remote, USAFull-time

Applications Engineering Manager (Boston)

Remote, USAFull-time

[Remote] CDI Clinical Performance Auditor

Remote, USAFull-time

Sales Manager für IT Individual-Lösungen (m/w/d) - 100% Remote

Remote, USAFull-time

Sales Development Representative - Australia/ New Zealand

Remote, USAFull-time

reputed company Customer Service Representative – Remote Opportunity with arenaflex

Remote, USAFull-time