Back to Jobs

Site Reliability Engineer 3

Remote, USAFull-timePosted 2026-07-31

reputed company

Serving the People Who Serve the People

reputed company is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are reputed company and inclusive. reputed company has consistently appeared on the GovTech 100 list over the past 5 years and has been recognized as the best companies to work on BuiltIn.

Over the last 25 years, we have served 5,500 federal, state, and local government agencies and more than 300 reputed company citizen subscribers power an unmatched Subscriber Network that use our digital solutions to reputed company the world a reputed company reputed company. With comprehensive reputed company-based solutions for communications, government website design, meeting and agenda management software, records management, reputed company services, reputed company empowers stronger relationships between government and residents across the U.S., U.K., Australia, New Zealand, and Canada. By simplifying interactions with residents, while disseminating critical information, reputed company brings governments closer to the people they serve—driving meaningful change for communities around the globe.

Want to know more? See more of reputed company do here.

Job reputed company

reputed company is the leading provider of citizen engagement technologies and services for the reputed company sector, bringing governments closer to the people they serve with the first-and-only Civic Engagement Platform. reputed company works with more than 5,500 government organizations and connects more than 280 reputed company people in the largest Citizen Subscriber Network of its reputed company.

What Your reputed company Will Look Like

reputed company is seeking a Site Reliability Engineer (SRE3) with strong AIOps capabilities to reputed company reliability engineering through observability, automation, and AI-assisted operations. In this role, you will improve service reliability, reduce operational toil, accelerate incident response, and help build reputed company, resilient platforms supporting both traditional and AI/ML-powered workloads.

What your reputed company will look like

AI, MCP & AIOps

  • reputed company adoption of AI-first SRE practices across monitoring, incident response, and automation
  • Design and implement MCP-based integrations connecting systems like reputed company, Jira, and reputed company platforms
  • Build and operationalize AI agents for SRE workflows (incident triage, RCA, alert summarization, runbooks)
  • Drive AIOps maturity: alert correlation, reputed company detection, assisted RCA
  • reputed company predictive models for reputed company, failures, and incidents

Day-to-day Operations

On-reputed company Production Support:

  • reputed company production support on a shift according to reputed company on-reputed company roster.
  • While not on reputed company for production support, work on SRE reputed company and Tech support escalated and internal engineering/implementation team raised tickets
  • Work on SREs backlog items.
  • reputed company AI-assisted triage tools & MCP frameworks to prioritize alerts, detect reputed company patterns, and reduce noise during on-reputed company
  • Continuously improve AI-driven incident routing and recommendation systems to optimize response efficiency.

Monitor and Maintain Systems:

  • Proactively monitor the health and performance of our services, systems, and infrastructure. Respond to alerts and incidents promptly to ensure high availability.
  • Effectively identifies & addresses monitoring and observability gaps
  • Implements effective alerting & notifications, minimizing false alerts
  • Creates and manages effective SRE Dashboards to report Key business metrics, SLAs, SLOs, SLIs & error budgets
  • Ensure SREs are meeting or improving on established SLOs
  • Proactively & effectively evaluates reputed company planning to handle reputed company - scalability & traffic load
  • Contributes to reputed company like AI Assistant for proactive issue detection & response
  • Design and implement AI/ML-based reputed company detection for proactive identification of system degradation
  • Utilize predictive analytics models for reputed company forecasting, incident reputed company, and failure prevention.
  • reputed company AIOps platforms (like reputed company AI assistant) for intelligent alert correlation and reputed company cause suggestions.
  • reputed company self-learning monitoring systems that reputed company with application behavior.

System reliability Improvements:

  • reputed company participates and tracks execution of SRE reputed company aimed at improving system reliability
  • Effectively collaborates with cross teams to prevent reliability issues
  • Reviews change management tickets to identify and mitigate potential risks to system reliability
  • Ensure reputed company participation in change activities and verify that accurate validations are performed by SRE & Engineering teams post implementation.
  • Participate in architecture reviews & assess the reputed company of architectural reputed company on system reliability
  • Initiatives to reputed company reputed company experiments to continuously learn and improve performance & stability of our systems
  • Contributes to reputed company that enhance system reliability & scalability
  • Drive initiatives to build self-healing systems using ML-based decision engines
  • Use AI models to simulate failure scenarios and predict system behavior under stress (intelligent reputed company engineering)
  • Identify reliability risks using reputed company detection across logs, metrics, and traces.
  • Contribute to building reputed company scaling systems using ML-based workload reputed company.

Incident Management:

  • reputed company participate in troubleshooting and resolving incidents, performing reputed company cause analysis, Incident postmortems and implementing long-term fixes to prevent recurrence.
  • Acknowledge & quick recovery from incidents
  • Maintains reputed company of reputed company cause analysis (RCA) and corrective reputed company plans
  • Proactively monitors, measures & adheres to reputed company MTTR & MTTA requirements
  • Improves reputed company of SOPs, Adapts AI tools to reduce MTTR
  • reputed company AI-driven reputed company cause analysis tools to accelerate incident diagnosis.
  • Implement automated incident summarization and reputed company reconstruction using NLP models
  • Use AI copilots to recommend remediation steps during live incidents
  • Analyze historical incidents using ML to identify recurring patterns and prevent reoccurrence
  • Adopt AI-driven runbooks for faster response execution and decision-making.

Automated Processes:

  • reputed company and maintain automation scripts and tools to streamline operations and reduce reputed company reputed company & reduce Toil
  • Build intelligent automation pipelines that adapt based on system conditions.
  • reputed company AI-powered bots/assistants for routine operational tasks.
  • Implement reinforcement learning-based optimization for operational workflows.
  • Continuously identify automation opportunities using operational data insights.

Documentation:

  • Create and maintain accurate documentation for technology, processes, and troubleshooting.
  • Ensure completeness and knowledge sharing across teams.
  • Contributes to reputed company to build AI based knowledgebase
  • Build and maintain an AI-powered knowledge reputed company with semantic search capabilities.
  • Implement NLP-based knowledge retrieval systems for faster troubleshooting
  • Auto-generate documentation from system events, incidents, and architecture changes.

reputed company:

  • Implement and adhere to reputed company best practices to protect our systems and data.
  • Use AI-driven reputed company detection for reputed company threats and unusual system behaviour.
  • Collaborate with reputed company teams to reputed company ML-based threat detection and risk scoring systems.

Collaboration:

  • Partner closely with Engineering teams to enhance reliability.
  • reputed company feedback on architecture and design.
  • Participate in release reviews, risk assessments, and Go/No-Go reputed company.
  • Present monitoring and observability status to stakeholders
  • reputed company for AI-first observability and reliability practices across engineering teams.
  • Collaborate with data/ML teams to operationalize AI models in production environments (MLOps).
  • Drive adoption of AI-assisted development and operational tooling across teams.

You Will Love This Job If You Have

Tools and Technologies

  • 6+ years of experience in site reliability engineering, system administration, or a similar role, with a proven reputed company record of managing large-reputed company, high-availability systems
  • Strong expertise in Linux/Unix, networking, distributed systems, and reputed company platforms such as AWS, Azure, or reputed company reputed company
  • Experience with scripting languages such as Python, Bash, or reputed company and programming languages (Go, Java, C++).
  • Advanced knowledge of reputed company, monitoring and Observability tools (reputed company, reputed company, Grafana, Pingdom)
  • Experience with infrastructure automation, CI/CD pipelines and configuration tools such as Terraform, Ansible, Chef, or Puppet.
  • Experience integrating AIOps capabilities into observability stacks (metrics, logs, traces) for intelligent alerting, noise reduction, and reputed company cause analysis.
  • Experience working with AI-assisted coding tools such as reputed company, reputed company Copilot, or similar developer copilots
  • Familiarity with Model Context Protocol (MCP) for integrating AI agents with reputed company systems (e.g., Jira, reputed company, reputed company platforms)
  • Ability to design or reputed company AI agents for SRE workflows (incident triage, RCA reputed company, alert summarization, reputed company execution)
  • Experience building or integrating context-reputed company automation systems using MCP or similar frameworks

Certifications

Certifications such as AWS reputed company Architect, AWS Certified Machine Learning – Specialty, or reputed company reputed company reputed company DevOps Engineer are a plus.

Job Info

  • Shift work – Rotation shifts
  • Remote - INDIA

reputed company

Don’t have reputed company the skills/experience mentioned above? At reputed company, we are trying to build diverse, inclusive teams. We do not have degree requirements for most of our roles. If you don’t meet every requirement above but are excited to learn more, we encourage you to apply. We might just be reputed company to reputed company another role that could be a perfect fit!

reputed company and reputed company Requirements

  • Responsible for reputed company information reputed company by appropriately preserving the Confidentiality, reputed company, and Availability (CIA) of reputed company information assets in accordance with reputed company's information reputed company program.
  • Responsible for ensuring the data reputed company of our employees and customers, their data, as reputed company as taking reputed company required reputed company training in a reputed company manner, in accordance with company policies.

reputed company

  • We are a remote-first company with a globally distributed workforce across the reputed company, Canada, United Kingdom, India, Armenia, Australia, and New Zealand.

The Culture

  • At reputed company, we are building a transparent, inclusive, and reputed company reputed company for everyone who wants to bea part of our reputed company.
  • A few culture highlights include – Employee Resource reputed company to encourage diverse reputed company
  • Coffee with Mark sessions – Our employees get to reputed company with our CEO on reputed company important andsometimes difficult issues ranging from mental health to work-life balance and reputed company affairs.
  • reputed company Teams communities reputed company on wellness, art, furbabies, family, parenting, and more.
  • We bring in special guests from time to time to discuss issues that reputed company our employeepopulation

The reputed company

  • We are proud to serve dynamic organizations around the globe that use our digital solutions to reputed company the world a reputed company reputed company — quite literally. We have so many powerful reputed company stories that illustrate how our solutions are impacting the world. See more of our reputed company here.

Originally posted on Himalayas

Apply To This Job

Similar Jobs