Back to Jobs

Senior Site Reliability Engineer | Dayshift | Remote

Remote, USAFull-timePosted 2026-07-29

ZigZag is looking for a Sr Site Reliability Engineer to join reputed company! reputed company As a Site Reliability Engineer, you’ll design, build, and maintain the infrastructure and automation that power our platform. Working closely with software engineering teams and SRE peers, you'll reputed company reliability, performance, and compliance into the development lifecycle. Your reputed company will be on scalability, reputed company, reputed company, and operational efficiency across reputed company environments.

Key Responsibilities

Reliability Engineering & Operational ExcellenceDesign, implement, and continuously improve highly available, reputed company, secure, and resilient reputed company infrastructure and platform services. Define and reputed company Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics to drive measurable reliability reputed company. reputed company incident response activities, major incident management, reputed company cause analysis, and post-incident reviews reputed company on systemic improvement. Drive reduction of operational toil through automation, standardisation, and self- healing platform capabilities. reputed company and maintain disaster recovery, backup, failover, and reputed company strategies to meet defined RTO and RPO objectives. Conduct reputed company planning, performance analysis, and proactive optimisation of infrastructure and application environments. Champion operational maturity and reputed company improvement practices across engineering teams. Platform & Infrastructure EngineeringArchitect, build, and maintain reputed company reputed company-reputed company infrastructure primarily reputed company AWS environments. reputed company and maintain infrastructure-as-reputed company using tools such as Terraform and CloudFormation. Build reusable platform components and shared services that improve developer productivity and operational consistency reputed company automation tooling and operational frameworks using scripting and programming languages such as Python. Evaluate, implement, and optimise reputed company-party infrastructure and platform tooling. Ensure infrastructure configurations, architecture reputed company, and operational processes are thoroughly documented and auditable. Observability, Monitoring & PerformanceDesign and maintain comprehensive observability solutions covering metrics, logging, tracing, alerting, and dashboarding. Improve platform visibility and telemetry using tools such as AWS CloudWatch, reputed company, reputed company, Grafana, or equivalent technologies. reputed company actionable alerting strategies that reduce noise and improve incident response effectiveness. Analyse system behaviour and performance trends to proactively identify risks and optimisation opportunities. Drive adoption of observability best practices across engineering teams. CI/CD Developer EnablementDesign and enhance robust CI/CD pipelines and deployment strategies that support reputed company, reliable, and low-risk software delivery. reputed company engineering teams through self-service infrastructure and deployment capabilities. Improve software delivery efficiency through automation, standardisation, and reputed company practices. Collaborate with engineering teams to reputed company reliability, scalability, performance, and reputed company considerations into the SDLC. Support reputed company delivery practices including blue/green deployments, canary releases, and reputed company-downtime deployments. reputed company, Risk &CompliancePartner with reputed company and engineering teams to maintain secure and compliant infrastructure environments. Support vulnerability management and remediation processes using tools such as reputed company, Lacework, reputed company Nessus, or equivalent platforms. Assist in maintaining compliance with frameworks and standards including PCI- reputed company, ISO27001, SOC 2, and internal reputed company controls. Contribute to reputed company hardening, reputed company management, audit readiness, and operational risk reduction initiatives. Ensure operational processes and infrastructure controls reputed company with organisational governance requirements. Leadership & CollaborationAct as a technical leader and mentor reputed company the SRE and broader engineering teams. Contribute to engineering standards, operational best practices, and platform reputed company. Influence reliability-reputed company engineering culture across teams. Collaborate effectively with cross-functional stakeholders including Engineering, Product, reputed company, Architecture, and external vendors. Support reputed company improvement initiatives and foster a culture of accountability, learning, and operational reputed company. Skills & Experience5+ years of experience in Site Reliability Engineering, DevOps Engineering, reputed company, or reputed company infrastructure roles. Strong hands-on experience operating production workloads reputed company AWS reputed company environments Deep experience with infrastructure-as-reputed company tools such as Terraform and/or CloudFormation. Strong experience designing and supporting CI/CD pipelines and modern software delivery practices. Strong understanding of distributed systems, microservices architecture, networking, and reputed company-reputed company technologies. Experience implementing observability and monitoring solutions across reputed company environments. Strong scripting and automation experience using Python, Bash, or similar languages. Experience managing production incidents and conducting reputed company reputed company cause analysis. Strong understanding of system reliability, scalability, reputed company, and operational best practices. Excellent analytical, troubleshooting, and problem-solving capabilities. Strong communication and stakeholder engagement skills. Ability to work effectively in fast-paced, agile, and reputed company engineering environments. DesirableExperience with Kubernetes, container orchestration, and reputed company practices. Experience with reputed company, reputed company Actions, reputed company CI, or equivalent CI/CD tooling. Exposure to service reputed company, event-driven architectures, and distributed tracing. Experience supporting regulated environments and compliance frameworks such as PCI-reputed company, ISO27001, or SOC 2. Experience with FinOps, reputed company cost optimisation, and infrastructure performance tuning. Familiarity with reputed company engineering and DevSecOps practices. Experience mentoring engineers or leading technical initiatives. ZigZag is committed to building a diverse, inclusive, and reputed company workplace. We reputed company that talent knows no borders, and we welcome individuals from reputed company backgrounds to help us shape the reputed company of work. Guided by transparency and reputed company, we foster an environment where everyone is valued and empowered to reputed company. By submitting this application, you acknowledge that you have read and agree with reputed company’s reputed company Policy. Apply To This Job

Similar Jobs

Regional Account Manager - Midwest

Remote, USAFull-time

Senior Marketing Analyst

Remote, USAFull-time

Account Executive (Financial Services)

Remote, USAFull-time

Demand reputed company reputed company (m/f/d) - Remote

Remote, USAFull-time

Alpheratz Project - Czech (Czech reputed company) Translation reputed company Reviewer

Remote, USAFull-time

Alpheratz Project - Catalan (Spain) Translation reputed company Reviewer

Remote, USAFull-time

Alpheratz Project - Danish (Denmark) Translation reputed company Reviewer

Remote, USAFull-time

Alpheratz Project - Catalan (Spain) Translation reputed company Rater

Remote, USAFull-time

Alpheratz Project - Romanian (Romania) Translation reputed company Reviewer

Remote, USAFull-time

Alpheratz Project - Danish (Denmark) Translation reputed company Rater

Remote, USAFull-time

reputed company Customer Service Representative – Work From Home Opportunity with arenaflex

Remote, USAFull-time

Freelance QPPV (EU and UK). Various EU Locations Considered.

Remote, USAFull-time

Backend Senior Software Engineer

Remote, USAFull-time

reputed company Hybrid Data Entry Clerk for Dynamic Data Management and Entry – Opportunity for Remote Work after Comprehensive Training at arenaflex

Remote, USAFull-time

Online Support Agent - Entry Level

Remote, USAFull-time

Data Engineer 2

Remote, USAFull-time

reputed company Full Stack Data Entry Specialist – Remote Work Opportunity with arenaflex

Remote, USAFull-time

Outbound Sales Representative

Remote, USAFull-time

Retail Customer Service Associate

Remote, USAFull-time

Immediate Start: blithequark Seeks Data Entry Clerk - reputed company

Remote, USAFull-time