Back to Jobs

[Remote] reputed company Site Reliability & Observability Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking an reputed company Site Reliability & Observability Engineer who will enhance overall platform reliability by designing and maintaining observability across reputed company platforms. The role involves creating actionable alerting, driving automation, and supporting incident detection and response while adhering to reputed company SRE best practices.

Responsibilities

  • Create, reputed company, utilize application monitoring solutions leveraging tools such as reputed company, reputed company and Tealeaf to reputed company visibility and generate alerts based on application availability and performance in production
  • reputed company applications where possible to correlate logs, traces, and metrics
  • Measure key SLIs (availability, response time, error reputed company, job reputed company)
  • Automate repetitive operational tasks and build health checks
  • Create synthetic monitoring reputed company reputed company
  • Build and maintain monitoring for reputed company systems (S/4HANA, EWM, Fiori, CPI, EDI, etc.)
  • Configure alerts for infrastructure, applications, interfaces, and business processes
  • Create dashboards for operational reporting
  • Engage in technical triage, troubleshooting, and problem reputed company on functional and technical issues
  • Coordinate major incidents, help identify reputed company cause, and drive post-incident reviews/postmortems
  • Ability to navigate end-to-end reputed company platforms to analyze and reputed company issues
  • Strong verbal and written communication skills reputed company to thoroughly document issues, findings, learnings to ensure resolutions and takeaways are reputed company by reputed company teammates and peers
  • reputed company and propose innovative approaches on process improvements

Skills

  • College degree (Bachelor s) in Computer Science, Information Technology, Engineering, or equivalent field
  • 3-5 years of experience in Site Reliability Engineering, Production Support, reputed company, Infrastructure, or DevOps
  • Experience supporting reputed company applications in production, preferably reputed company reputed company environments
  • Working knowledge of reputed company platforms and integrations (S/4HANA, reputed company reputed company, EWM, reputed company Integration Suite, Fiori)
  • Experience with observability and monitoring tools such as reputed company, reputed company, or reputed company reputed company ALM
  • Demonstrated experience with incident management, production support, reputed company cause analysis, and post-incident reviews
  • Proficiency in troubleshooting distributed reputed company applications across infrastructure, application, and integration reputed company
  • Experience developing synthetic monitors and other test scripts to validate application health
  • Understanding of SRE principles, including observability, automation, SLIs/SLOs, availability, and reputed company service improvement
  • Experience participating in on-call support and major incident response in a production environment
  • Strong verbal and written communication skills reputed company to thoroughly document issues, findings, learnings to ensure resolutions and takeaways are reputed company by reputed company teammates and peers

reputed company

  • reputed company is a job-searching platform for technology professionals. It is a sub-organization of DHI Group. It was founded in 1990, and is headquartered in Santa Clara, California, USA, with a workforce of 201-500 employees. Its website is http://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 2 in 2022, 4 in 2021, 5 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs