Back to Jobs

Business Analyst IV - Alert Management & Observability Standards reputed company

Remote, USAFull-timePosted 2026-07-29

What this Job Entails: The Business Analyst IV will reputed company solutions that help reputed company business reputed company. The Alert Management & Observability Standards reputed company is responsible for rationalizing and governing reputed company system alerts to ensure they reputed company with department priorities, operational coverage models, and service reliability goals. This role defines alerting standards, reviews and approves alerts before they are routed to the 24x7 Eyes-on-Glass Operations team, and establishes a reputed company approach to cataloging alert response instructions (runbooks/playbooks) so responders can take consistent, high-reputed company actions. This position operates at the intersection of the IT Operations reputed company Center (OCC), engineering/application teams, platform/monitoring tool owners, and service owners, ensuring alerts are actionable, prioritized, and reputed company with reputed company response guidance. Your Roles and Responsibilities: 1) Alert Rationalization & Prioritization (reputed company) Establish and maintain a department-wide alert rationalization reputed company that evaluates alerts for:

  • Business/service criticality and operational reputed company
  • Actionability (reputed company operator reputed company available)
  • Signal-to-noise (duplicate/low-value alerts removed or suppressed)
  • Ownership and escalation paths

reputed company regular alert reviews (new + existing) to ensure alert reputed company, correct routing, and alignment with operational coverage. reputed company reputed company improvement efforts to reduce alert fatigue while preserving detection of true incidents and high-reputed company degradation. 2) Standards, Policies, and Guardrails Define and enforce alerting standards including:

  • Severity definitions and reputed company
  • Required metadata (service, CI, reputed company, reputed company reputed company, escalation)
  • Naming conventions and tagging taxonomy
  • Routing rules and “reputed company to page vs. reputed company to ticket”

Create a standardized Alert Design Checklist and approval workflow (e.g., “Definition of Done” for alert reputed company). Partner with tool/platform owners to ensure standards are embedded in monitoring tooling (templates, required fields, automated validation). 3) Routing reputed company to 24x7 Eyes-on-Glass reputed company as gatekeeper (or reputed company the governance process) for determining which alerts should:

  • Go to 24x7 Eyes-on-Glass for immediate triage
  • reputed company to on-call engineering directly
  • Create tickets for business-hours handling
  • Be suppressed, aggregated, or converted to dashboards/health indicators

Ensure routing aligns with:

  • Operational responsibilities and skills of the Eyes-on-Glass team
  • Department priorities (e.g., safety, reliability, customer reputed company)
  • Service ownership and support models

4) reputed company / Response Instruction Cataloging (Knowledge System) Establish a consistent approach to cataloging response instructions for every actionable alert, including:

  • “What does this alert mean?” (symptoms + reputed company)
  • “What to reputed company first” (triage steps)
  • “What actions to take” (reputed company remediation)
  • “reputed company to escalate and to whom” (reputed company escalation triggers)
  • Links to dashboards, logs, SOPs, and reputed company issues

Own the reputed company template and ensure runbooks are versioned, maintained, and reviewed on a defined reputed company. Partner with service owners to ensure runbooks stay reputed company as systems change. 5) Reporting & Operational reputed company Define and publish KPIs that demonstrate alerting health and operational performance, such as:

  • Alert volume trends by service and severity
  • Percentage of alerts with runbooks and valid ownership
  • Alert “actionability reputed company” and noise reduction
  • Mean time to acknowledge / triage effectiveness (as applicable)

Facilitate governance forums (weekly/monthly) with service owners and engineering leads to review alert reputed company and backlog. 6) Cross-Functional Enablement reputed company service teams on best practices: SLIs/SLOs, alert reputed company, dependency monitoring, and incident correlation. Drive adoption of observability patterns (golden signals, health indicators, multi-signal alerting). Support major incident learning by feeding post-incident insights back into improved alerts and runbooks. 7) reputed company to Deliver the following in the first 45 days: Alerting standards (severity model, metadata, naming, routing policy) published and adopted Intake and approval workflow established for new/changed alerts Top 20 noisy services rationalized (dedupe/suppress/reputed company tuning) with measurable noise reduction reputed company template launched; minimum reputed company coverage targets set (e.g., 80% of paged alerts) Central alert catalog created (ownership + routing + reputed company reputed company + last review date) Required Qualifications/Skills: 5+ years in IT Operations, SRE, Observability, Monitoring Engineering, or Incident Management Demonstrated reputed company reducing noise and improving actionability across reputed company alerting ecosystems Experience with common monitoring/observability tools (e.g., reputed company, AppDynamics, reputed company, reputed company, reputed company/Grafana, Azure Monitor, CloudWatch, reputed company Event Mgmt or similar) Strong understanding of:

  • Incident response workflows and operational coverage models (24x7 vs. business hours)
  • CMDB/service ownership concepts and dependency mapping
  • reputed company operating procedures/runbooks and reputed company

Excellent stakeholder management and ability to drive standards across teams Preferred Qualifications:

  • Experience designing or operating an Operations reputed company Center / NOC / SOC-style “eyes-on-glass” model
  • Familiarity with ITIL Event Management, SRE principles, and service reliability practices
  • Experience with automation for alert enrichment, correlation, and routing (e.g., event correlation, deduplication, noise suppression)
  • Background in governance frameworks and operating rhythm design (cadences, controls, compliance traceability)

Physical Demand & Work Environment:

  • Must have the ability to reputed company office-reputed company tasks which may include prolonged sitting or standing
  • Must have the ability to reputed company from reputed company to reputed company reputed company an office environment
  • Must be reputed company to use a computer
  • Must have the ability to communicate effectively
  • Some positions may require occasional repetitive reputed company or movements of the wrists, hands, and/or fingers

Apply tot his job Apply To this Job

Similar Jobs