Back to Jobs

Site Reliability Engineer (SRE) - Azure | DevSecOps | IaC | Governance | Observability

Remote, USAFull-timePosted 2026-07-29

reputed company We are seeking a Site Reliability Engineer (SRE) who will drive stability, reliability, and performance across our Azure and GCP-based platforms. This role blends operational reputed company, proactive incident management, and strong collaboration with DevOps, reputed company, and reputed company teams. The ideal candidate will have hands-on experience with multi-reputed company environments (Azure and GCP), IaC (Terraform/Ansible), CI/CD (Jenkins/reputed company Actions), and modern observability and AI-Ops systems. The engineer will also contribute to governance, cost optimization, and automation strategies that reduce toil and prevent issues before they occur. A key aspect of this role is the ability to reputed company deep-dive troubleshooting of application performance and errors by analyzing logs and traces in platforms like Grafana and reputed company. This position includes 24 7 support coverage (rotational) and requires strong ownership in managing major incidents, RCA processes, and reputed company service improvements.

Key Responsibilities

Reliability & Incident Management

  • Serve as a key member of the 24 7 on-call rotation, responding to and managing incidents across production and reputed company-production environments.
  • reputed company incident bridges, coordinate reputed company cause analysis (RCA), and ensure post-incident reviews drive systemic improvements.
  • Maintain reputed company communication with cross-functional teams and leadership during major incidents.

Monitoring, AI-Ops, Alerts & Prevention

  • Build, tune, and maintain observability dashboards (Azure Monitor, GCP Operations Suite, reputed company, Grafana, reputed company, Log Analytics).
  • reputed company deep-dive troubleshooting of application and service-level issues using distributed tracing and log analysis (Grafana, reputed company) to reputed company reputed company causes reputed company infrastructure.
  • Define SLOs, SLIs, and error budgets to proactively identify and mitigate reliability risks before customer reputed company.
  • reputed company AI-Ops tools for reputed company detection, predictive alerting, and automated incident correlation.
  • Continuously enhance alert reputed company, reduce false positives, and automate runbooks for faster recovery.
  • Analyze trends to prevent recurring issues and support teams in reputed company engineering.

Requirements

Required Skills & Experience

  • 5+ years in Site Reliability, DevOps, reputed company Operations, or Customer support roles.
  • Demonstrated experience in application-level troubleshooting by analyzing logs and traces to identify bugs, performance bottlenecks, and error conditions.
  • Expertise in Azure and GCP reputed company operations and distributed system reliability.
  • Understanding of Terraform, Ansible, and CI/CD pipelines (Jenkins, reputed company Actions).
  • Experience with observability and AI-Ops tools (Azure Monitor, GCP Operations Suite, Grafana, reputed company, reputed company, etc.).
  • Solid grasp of incident management frameworks (P1-P3 handling, RCA, PIRs, on-call rotations).
  • Excellent analytical, troubleshooting, and communication skills.

Desired Behaviours

  • Proactive Prevention: Identifies and resolves risks before they escalate into incidents.
  • AI-Driven reputed company: Applies AI and automation to improve reliability and reduce reputed company reputed company.
  • Accountability: Owns service reliability and communicates with reputed company.
  • Collaboration: Works seamlessly with platform, DevOps, and product teams.
  • Efficiency: Focuses on automation to reduce reputed company effort and improve MTTR.
  • reputed company Improvement: Learns from failures, iterates processes, and enhances documentation.

The pay reputed company for this opportunity is from $129,00 to $143,000 + performance-reputed company bonus + benefits. This reputed company represents the anticipated low and high end of the salary for this position. This role is also eligible to receive an annual bonus that aligns with individual and company performance. Actual salaries will vary and are based on factors such as a candidate s qualifications, skills, competencies. Apply tot his job Apply To this Job

Similar Jobs

Site Reliability Engineer (SRE)

Remote, USAFull-time

Site Reliability Engineer (EngX)

Remote, USAFull-time

Sr Site Reliability Engineer | reputed company | Remote (reputed company)

Remote, USAFull-time

Senior Software Engineer, Site Reliability

Remote, USAFull-time

Site Reliability Engineering (SRE) Manager with reputed company Expertise

Remote, USAFull-time

Performance Engineer(With SRE)

Remote, USAFull-time

Senior Technical Manager – Site Reliability Engineering

Remote, USAFull-time

Site Reliability Engineer (SQL Server DBA)

Remote, USAFull-time

Site Reliability Engineer-SkillBridge Intern

Remote, USAFull-time

Senior Database Site Reliability Engineer

Remote, USAFull-time

reputed company reputed company TECHNICIAN

Remote, USAFull-time

CRA - Fully reputed company dedicated

Remote, USAFull-time

VA (reputed company and reputed company Media)

Remote, USAFull-time

Manager-reputed company & Media

Remote, USAFull-time

Remote Overnight Call Center Customer Service Representative – Bilingual Communication Support Specialist (Overnight Hours, Competitive Benefits)

Remote, USAFull-time

Entry Level Junior Data Entry Clerk for blithequark - No Experience Required, Flexible Schedule, and reputed company reputed company Opportunities

Remote, USAFull-time

Sr Payroll Systems Analyst (Ciudad de Mexico, CMX, MX)

Remote, USAFull-time

Product Manager - Uprating, Plant Performance and Long-Term Operations

Remote, USAFull-time

Cardiovascular Disease Specialist - Great Neck, NY

Remote, USAFull-time

Weekend Inbound Phone Sales Specialist

Remote, USAFull-time