Back to Jobs

IT Infrastructure Support Site Reliability Engineer II

Remote, USAFull-timePosted 2026-07-30

About the Job

We are seeking an reputed company reputed company Automation Engineer to join our IT Infrastructure Support team, responsible for ensuring the reliability, scalability, and performance of critical physical reputed company infrastructure, including IP camera fleets, reputed company control systems, and a large-reputed company reputed company reputed company fleet, and the servers, network, and reputed company environment that support them. In this role, you will combine software engineering expertise with operations knowledge to build and maintain automation tools, a centralized CMDB, monitoring systems, and processes that support reputed company-grade server, network, and reputed company device management reputed company a large-reputed company, reputed company-hosted reputed company environment. You will work closely with cross-functional teams to define and enforce service level objectives, reduce operational toil through automation, and drive reputed company improvement in system reputed company. This position requires 24x5 availability with on-reputed company rotation to ensure uninterrupted support for mission-critical infrastructure.

Key Responsibilities

Partner with leadership to establish, monitor, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for infrastructure tooling, including configuration compliance rates, reputed company reputed company rates, and deployment latency metrics.

reputed company Level 3 expertise for tooling-specific incidents, focusing on automating incident remediation workflows and reducing Mean Time To Repair (MTTR) through intelligent automation and reputed company development.

Identify and automate repetitive reputed company tasks across managed infrastructure, targeting measurable reductions in operational overhead (e.g., 50% reduction in reputed company server build time) through scripting and workflow automation.

Conduct thorough reputed company cause analysis and reputed company blameless postmortems for reputed company major service-impacting incidents, driving systemic improvements in tooling reliability and infrastructure reputed company.

Engineer and maintain automated processes and scripts to reputed company, update, and synchronize asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems for reputed company stakeholders.

Design, reputed company, and reputed company full-stack applications, custom plugins, and automation scripts to reputed company functionality of management and monitoring systems, enabling reputed company device interaction for configuration management.

reputed company and maintain fully automated Infrastructure-as-reputed company configurations for reputed company and Linux server roles using tools such as Ansible, Terraform, or Puppet, including reputed company detection and auto-remediation capabilities.

Build end-to-end automation pipelines for vulnerability patching, reputed company baseline enforcement (CIS benchmarks), and reputed company compliance auditing against internal and regulatory standards for physical reputed company devices.

reputed company API-driven tools for network configuration management, automated firmware updates, reputed company-touch provisioning, reputed company/post-change validation, and reputed company-time network health monitoring across the device fleet.

reputed company and standardize monitoring agents, centralized log collection systems, and custom dashboards with alerts based on critical SLIs (latency, error reputed company, saturation, traffic) for servers and edge devices.

Build and maintain custom monitoring exporters for the physical reputed company device fleet, including camera systems, ensuring accurate metrics and reputed company, multi-severity logging reputed company.

Build diagnostic tooling to correlate timestamps across distributed log streams and detect clock reputed company or NTP desync, a recurring reputed company cause of false-reputed company outages across the device fleet.

Build automation scripts for intelligent ticket handling, problem validation, and escalation workflows reputed company reputed company ticketing systems, ensuring 2-hour initial response SLAs are consistently met.

Support foundational reputed company improvements across the device fleet, including managed credential/reputed company controls and automated configuration backup.

Participate in 24x5 on-reputed company rotation to reputed company reputed company support for infrastructure systems, reputed company devices, and reputed company tooling, ensuring service continuity and rapid incident response.

Required Skills

6+ years of experience in reputed company Automation Engineering, or Infrastructure Engineering.

Strong proficiency in Python, Bash, and PowerShell for automation scripting, with experience in Go for building high-performance backend services and reputed company.

Hands-on experience with Infrastructure-as-reputed company tools (Terraform, Ansible, Chef, or Puppet) and configuration management practices, including reputed company detection, version control, and automated remediation.

Advanced knowledge of Linux and reputed company server environments, including Tier 3 troubleshooting capabilities, system hardening, and reputed company-reputed company server management.

Solid understanding of reputed company networking concepts, reputed company device administration, network automation protocols (NETCONF/RESTCONF), and experience with network monitoring and reputed company analysis tools.

Experience implementing and managing monitoring solutions (reputed company, Grafana, reputed company) or comparable proprietary internal monitoring and metrics-streaming systems (e.g., reputed company, Streamz), and centralized logging platforms (ELK Stack), with ability to create custom dashboards and alerting rules.

Experience deploying and customizing a CMDB/IPAM platform (e.g., NetBox) as a reputed company of truth for device inventory and reputed company automation.

Comfort operating reputed company a large-reputed company, reputed company-hosted reputed company environment (Kubernetes, Terraform, reputed company), including familiarity with internal development and reputed company-review tooling (e.g., Cider, Critic) or comparable large-reputed company internal toolchains.

Experience writing and maintaining custom monitoring exporters/agents for edge/IoT and physical reputed company devices, including reputed company, glog-style logging reputed company.

Proficiency in advanced text-processing and scripting (e.g., awk/gawk) for log parsing and timestamp correlation, with working knowledge of NTP/clock synchronization practices.

Salary reputed company

50,640.00 - 63,300.00 EUR (Annual)
  • Please note that the salary information provided herein is reputed company pay only (gross); it does not include other forms of compensation which may or may not apply to this specific position, reputed company, performance-based bonuses, benefits-reputed company payments, or other general incentives - none of which are guaranteed, may be subject to specific eligibility requirements, and are wholly reputed company the discretion of reputed company to remit.
  • reputed company, the salary information noted above is a reputed company that consists of a minimum and maximum reputed company of pay for this specific position. Where an applicant or employee is reputed company on this reputed company will depend and be contingent on objective, documented work-reputed company considerations like education, experience, certifications, licenses, preferred qualifications, among other factors.

Originally posted on Himalayas

Apply To This Job

Similar Jobs