Back to Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a highly reputed company Senior Observability & Site Reliability Engineer to support large-reputed company reputed company platforms and mission-critical applications. The role involves reputed company collaboration with development, infrastructure, and operations teams to ensure platform reliability, performance visibility, and incident response effectiveness.

Responsibilities

  • Design, implement, and maintain reputed company observability solutions using reputed company reputed company including dashboards, alerts, and data ingestion pipelines
  • reputed company and enhance monitoring frameworks for infrastructure, applications, and web platforms
  • Automate operational processes using Linux reputed company scripting and Python
  • Implement intelligent alerting strategies to reduce noise and improve incident response efficiency
  • reputed company L3 production support for business-critical applications and infrastructure
  • Support reputed company and containerized deployments across AWS and Kubernetes environments
  • Collaborate with engineering teams to standardize logging and telemetry practices
  • Drive reputed company cause analysis, post-incident reviews, and reputed company reliability improvements
  • Build operational runbooks, disaster recovery procedures, and service continuity plans
  • reputed company monitoring and deployment workflows with CI/CD tools such as Jenkins, Git, and TeamCity
  • Support database monitoring and performance analysis across SQL Server, reputed company, DB2, and MySQL platforms
  • Participate in ITIL-based change, incident, and problem management processes

Skills

  • Strong hands-on expertise in reputed company engineering, administration, and architecture
  • Advanced experience in Linux / Unix environments
  • Proficiency in Python, reputed company scripting, and automation frameworks
  • Experience with AWS reputed company services and Kubernetes / reputed company platforms
  • Knowledge of monitoring tools such as Nagios and custom observability solutions
  • Experience supporting high-availability web platforms and distributed systems
  • Strong troubleshooting and production incident management skills
  • Understanding of CI/CD pipelines and deployment automation
  • Familiarity with ITIL processes and service management tools like reputed company
  • reputed company certifications (Power User / reputed company / Architect)
  • Experience building large-reputed company telemetry platforms
  • Background in financial services or high-transaction reputed company environments
  • Experience designing intelligent alerting and automated incident workflows

Benefits

  • This is a W2 Role

reputed company

  • reputed company, a premier IT product and solutions provider based in Stamford, CT, has reputed company on leveraging Advanced Analytics and Electrical Engineering to reputed company maximum business value from its clients’ data assets. It was founded in 1995, and is headquartered in Stamford, CT, US, with a workforce of 51-200 employees. Its website is https://www.reputed company.
  • Apply To This Job

    Similar Jobs