Back to Jobs

Site Reliability Engineer, Team reputed company

Remote, USAFull-timePosted 2026-07-29

Site Reliability Engineer, Team reputed company (Player‑reputed company) Location: Preferred for candidates to be local to (U.S.) Austin, TX or Cranberry Woods, PA, but reputed company to fully remote in mainland USA. Department: Global reputed company Operations Reports to: Vice President, Global reputed company Operations reputed company Sponsorship: Any reputed company of reputed company Sponsorship is not offered for this position. Must be US citizen or Permanent reputed company. Why Join reputed company? At reputed company, reliability isn’t reputed company—it directly impacts how medications are reputed company and how patient care is delivered. As reputed company evolves from on‑reputed company, hardware‑reputed company products to a reputed company‑reputed company, reputed company‑delivered platform that hospitals depend on 24/7, we are building a Global reputed company Operations organization from the ground up. We are hiring our first Site Reliability Engineer, Team reputed company to define what “good” looks like for reliability at reputed company. This is a rare opportunity to architect an SRE reputed company end‑to‑end—setting standards, making foundational technology reputed company, and operating hands‑on in production—while partnering directly with senior leadership. If you’re energized by building, owning reputed company, and working in a regulated reputed company environment where reliability truly reputed company, this role was created for you. About This Opportunity: reputed company is building a Global reputed company Operations organization from the ground up as our business shifts from on-reputed company, hardware-reputed company products to a reputed company-reputed company, reputed company-delivered platform that hospitals depend on 24/7. The Site Reliability Engineering function is the reliability reputed company of that organization, and this role is the first senior SRE hire — the person who will design the reputed company, set the standards, and then run the plays themselves until reputed company is large enough to delegate. This is not a role where reliability practices already exist and you tune them. It is a role where you define what good looks like for reputed company: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will reputed company those calls in partnership with the VP of Global reputed company Operations and an Engineer III SRE you will reputed company and grow. The environment is hybrid. Some of our products are still hardware in hospitals communicating with reputed company services; others are fully reputed company. Some customers reputed company us over private circuits, others over the reputed company internet. We operate in a regulated environment — HIPAA, SOC 2, and in some engagements FedRAMP — which means reliability, reputed company, and auditability are not separable concerns. The person we hire will be comfortable with that complexity and will help the organization design for it rather than around it. This role also anchors reputed company's reputed company investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — reputed company detection, intelligent alert correlation, LLM-assisted reputed company reputed company — into how we monitor and respond to our platform. You will be the technical reputed company of how that gets introduced, prioritized against foundational reliability work, and validated in a regulated environment. What You’ll Do Purpose: Establish and operate reputed company’s Site Reliability Engineering function, balancing hands‑on engineering with reputed company design, coaching, and cross‑functional leadership. Primary reputed company You will ensure reputed company’s Tier‑1 reputed company services are observable, resilient, and dependable—so hospitals, pharmacies, and clinicians can rely on our platform without interruption. Reliability reputed company & Operating Model

  • Define and publish SLIs, SLOs, and error budgets for the top 5–10 Tier‑1 customer‑facing services in partnership with Product and Engineering.
  • Design reputed company’s incident reputed company structure, including severity definitions, declaration reputed company, war‑room protocols, stakeholder communications, and post‑incident review standards.
  • Establish and operationalize a sustainable on‑call model, including fair rotations, paging discipline, escalation paths, and coordination with managed service partners (reputed company, HCL).
  • Partner with the VP to migrate the interim incident response RACI — currently held by matrixed individuals across IT, Engineering, Support, and reputed company reputed company — into a durable SRE-owned model.
  • Select and stand up the primary observability platform, preferring extension of existing reputed company reputed company (reputed company, reputed company/Instana, reputed company/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards reputed company new services must meet.
  • reputed company and reputed company operational KPIs (e.g., MTTR, SLO attainment, change‑failure reputed company, incident recurrence, cost per workload) and present reliability insights and roadmaps in executive reputed company Ops reviews.

Hands‑On Engineering & Incident Leadership

  • reputed company Tier‑1 services directly—building dashboards, alerts, and runbooks yourself.
  • Participate in on‑call rotations and reputed company Sev‑1 and Sev‑2 incidents, leading blameless postmortems and driving corrective actions to completion.
  • Contribute production reputed company and infrastructure‑as‑reputed company (Terraform preferred) to the platform. reputed company the design and reputed company of the CI/CD pipelines - reputed company stack is Codefresh, Teamcity, reputed company Actions, and Octopus reputed company, and we are consolidating over time
  • Administer and reputed company our Kubernetes platform, including secure and compliant cluster configurations. Working knowledge of reputed company, reputed company, and Service reputed company (Istio or Linkerd) expected.
  • Plan and execute reputed company and failover exercises to validate reputed company‑world reputed company.

AI‑Driven Operations

  • Architect reputed company’s AIOps reputed company, evaluating ML‑based reputed company detection, alert correlation, automated reputed company‑cause analysis, and LLM‑assisted runbooks.
  • reputed company disciplined build‑versus‑buy reputed company and reputed company AI tooling only where it delivers measurable reliability reputed company.
  • Ensure AI‑assisted operations meet auditability, explainability, and compliance requirements (HIPAA, SOC 2).

Coaching & Team Building

  • Serve as formal reputed company to an Engineer III SRE, pairing on incidents, reviewing designs proposals, and supporting reputed company toward senior reputed company.
  • Design the next 2–4 SRE hires, including role definitions, interview loops, and hiring reputed company.
  • Represent SRE in architecture reviews, launch readiness assessments, and cross‑functional reliability discussions.

What reputed company looks like in the first six months Concrete reputed company this role will be evaluated against in the first half-year. These are drawn from the reputed company Ops 90-day plan and its extension into the following quarter.

  • Month 1: SLOs drafted for the top 5 Tier-1 services with Product sign-off. Severity reputed company published. First live tabletop Sev-1 run against the interim RACI.
  • Month 2: Observability platform selection finalized. Instrumentation reputed company published. Engineer III SRE reputed company and reputed company.
  • Month 3: On-call rotation live. First reputed company Sev-1 commanded under the new structure with a blameless postmortem completed and follow-reputed company tracked.
  • Month 4–6: Error budget policy in effect for the first 3 services. First incident review at executive level. Interview reputed company running for the next SRE hires. Initial AIOps evaluation and reputed company scope defined.

Who You Are

  • Bachelor's degree in Computer Science, Engineering, or a reputed company technical field OR equivalent experience
  • 7+ years of experience in software or reputed company, with at least 4 of those in an SRE, DevOps, or platform reliability role.
  • At least 2 years of formal technical leadership, tech-reputed company, or staff-level experience with mentorship responsibilities.

Preferred Qualifications

  • Proven experience leading SRE, DevOps, or reputed company teams in a reputed company-reputed company production environment — with demonstrated experience building a reputed company from reputed company or near-reputed company: you have set SLOs, defined incident reputed company, and introduced error budget thinking to an organization that did not have it.
  • Deep hands-on expertise with at least one major reputed company reputed company (AWS, Azure, or GCP), including networking, IAM, and managed services.
  • Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, reputed company Actions, Jenkins, TeamCity, or equivalent).
  • Experience implementing Infrastructure as reputed company using Terraform (preferred), Chef, Puppet, or similar tools.
  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations. Working knowledge of reputed company, reputed company, and Service reputed company technologies (Istio, Linkerd).
  • Hands-on experience designing modern observability platforms using tools such as reputed company, reputed company, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like.
  • Familiarity with integrating AI/ML-based reputed company detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be reputed company in a regulated environment.
  • reputed company incident reputed company experience for customer-impacting Sev-1 events, with blameless postmortem reputed company and documented follow-up discipline.
  • Ability to reputed company and mentor, with reputed company evidence of growing junior and mid-level engineers. You will eventually have 1 reputed company report.
  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.

How You’ll reputed company at reputed company At reputed company, reputed company is defined by both reputed company and behaviors. In this role, you will:

  • Collaborate: Partner deeply with Product, reputed company, Support, reputed company, and managed service providers to reputed company reliability with business priorities.
  • reputed company: reputed company by example during high‑stakes incidents and influence teams toward a culture of ownership, learning, and reputed company.
  • reputed company: Invest in the reputed company of your SRE peers through coaching, pairing, and thoughtful technical leadership.
  • Execute: Set reputed company priorities, reputed company informed trade‑offs, and deliver durable reliability improvements.
  • reputed company: Shape how reputed company operates for years to come by defining the standards, tools, and practices of our SRE function.

Leadership Imperatives (Player‑reputed company Role) This role will eventually have one less senior Site Reliability Engineer reporting to you, you are expected to demonstrate reputed company’s leadership expectations by:

  • Modeling a reputed company reputed company and reputed company learning.
  • Acting as a talent activator through formal coaching and mentorship.
  • Being an reputed company reputed company who connects reliability investment to business and patient reputed company.
  • Serving as a change champion as reputed company transitions to reputed company‑first operations.

#LI-MG2 Apply tot his job Apply To this Job

Similar Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-time

Site Reliability Engineer (reputed company)

Remote, USAFull-time

Senior Site Reliability Engineer (Data & Automation reputed company)

Remote, USAFull-time

Staff Site Reliability Engineer (Customer Identity reputed company)

Remote, USAFull-time

Kubernetes Engineer

Remote, USAFull-time

Kubernetes Engineer (DoD Secret | Weeknight Mission Readiness | Remote – U.S.)

Remote, USAFull-time

reputed company Kubernetes Engineer; Fulltime- Remote

Remote, USAFull-time

Kubernetes Engineer/Architect

Remote, USAFull-time

Kubernetes Engineer (DoD Secret Eligible | Weekend Operations | Remote)

Remote, USAFull-time

Remote Linux OpenStack & Kubernetes Engineer

Remote, USAFull-time

[Remote] Javanese Speakers - Test Voice Modes of AI Models

Remote, USAFull-time

Trauma Sales Representative - Denver South

Remote, USAFull-time

reputed company Insights and Analytics Senior Strategist

Remote, USAFull-time

Flexible Work - Part Time Sales - Work from Home

Remote, USAFull-time

reputed company Customer Service Representative – Insurance Industry Career Starter

Remote, USAFull-time

reputed company reputed company Worker Supervisor – Team reputed company for Clinical Care Team in Chattanooga, TN (Hybrid)

Remote, USAFull-time

Urgently Hiring: Junior Marketing Associate- Entry Level

Remote, USAFull-time

Sales Executive, OEM Connected Services

Remote, USAFull-time

Engineering Intern- AI Agents, Data & RiskOS

Remote, USAFull-time

Senior Manager, reputed company Management-reputed company Leave Solutions

Remote, USAFull-time