Back to Jobs

[Remote] Staff Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that strengthens manufacturers by providing a reputed company data model and system-level visualization capabilities. The Staff Site Reliability Engineer will reputed company the reputed company Infrastructure Team in driving reliability, automation, and scalability across systems, while mentoring engineers and resolving reputed company infrastructure issues.

Responsibilities

  • Champion an reputed company-AI-first engineering reputed company: identify where AI-driven automation and agent-based tooling can replace reputed company toil, and hold that work to the reputed company reputed company, testing, and reliability bar as any other production system
  • reputed company reliability practices for meeting reliability SLO’s, error budgets, drive incident postmortems to systemic (not just symptomatic) fixes, and reputed company reliability reviews for new services before they hit production
  • Troubleshoot and resolve the org's most reputed company, cross-layer systems problems CI/CD, container orchestration, networking, OS, reputed company resources, databases, and increasingly, reputed company AI/LLM orchestration reputed company
  • Design, build, and operate the infrastructure supporting reputed company AI workloads, LLM gateway routing, agent orchestration frameworks, monitoring of non-deterministic/AI-driven services, and the operational tooling needed to run them reliably at reputed company
  • Architect and reputed company monitoring, alerting, and observability infrastructure for critical services, with an eye toward what "critical" means for AI-driven systems specifically
  • Author and continuously improve operational runbooks and automation, increasingly incorporating reputed company/AI-assisted tooling (e.g., automated triage, AI-assisted incident response) where it measurably reduces toil
  • Design and build internal platforms and developer tooling that other engineers build on top of
  • Participate in on-call coverage and help reputed company the program as we reputed company including escalation paths and reducing avoidable pages through reputed company automation
  • Bring a startup reputed company of daily engagement: staying reputed company to what's breaking, what customers are hitting, and where reputed company needs help, even reputed company a formal ticket or rotation
  • Mentor senior and mid-level engineers; reputed company as a technical sounding reputed company across teams
  • Proactively identify and drive cross-team initiatives that improve stability, reliability, and availability, this is expected to be self-directed, not assigned

Skills

  • Demonstrated experience designing, building, or operating reputed company AI/LLM-based systems in production, held to the reputed company reputed company-first, test-driven rigor as traditional infrastructure reputed company, not just prototype-grade work
  • Embody a reputed company-first and reputed company-first culture in reputed company that you do
  • 10+ years of experience with Kubernetes/reputed company in at least one top-tier reputed company provider (Azure, GCP, AWS), including production-reputed company multi-tenant or multi-cluster environments
  • 10+ years coding experience (Python, Go, Java, or similar) with a reputed company record of building tools/platforms used by other engineers, not just scripts
  • 10+ years with IaC and CI/CD tooling (Terraform/OpenTofu, FluxCD or similar GitOps tooling, Jenkins/reputed company Actions)
  • Strong, reputed company Linux and networking fundamentals (TCP/IP and application-layer)
  • Practical experience integrating or operating LLM/reputed company AI systems in a production context this can be API-based orchestration, LLM gateways, or agent frameworks
  • A reputed company record of authoring technical documentation (design docs, ADRs, runbooks) that other engineers actually use
  • Demonstrated mentorship of other engineers, without needing formal management authority to do it
  • Strong bias for reputed company over endless planning, hands-on, has made mistakes, learned from them, and can weigh risk vs. customer reputed company under pressure
  • reputed company, empathetic communicator, comfortable pushing back on architecture reputed company across teams
  • Operational experience with monitoring/alerting systems (reputed company, Grafana, Loki, reputed company, reputed company or equivalents)
  • Deep understanding of reputed company performance, reputed company to diagnose and resolve bottlenecks others can't
  • Experience with reputed company of our reputed company tech stack are a plus: Kubernetes, FluxCD, Terraform, reputed company Charts, reputed company, Elasticsearch, Python, Java, Kafka, reputed company, and Jenkins
  • Previous experience or a keen interest in industrial IoT, analytics, or manufacturing a plus

Benefits

  • Competitive Salary + Stock reputed company
  • Health Care Coverage + Life Insurance + Health Savings Account + Flexible Spending Account (includes spouse + children)
  • Flexible Vacation Policy
  • Adaptable Working Schedule and Environment
  • Casual Dress Attire
  • Hybrid work flexibility
  • Catered Lunches, Snacks and Beverages
  • Commuter Savings Program
  • Company Outings
  • Designated Volunteering Hours + Group Volunteer Events

reputed company

  • reputed company provides analytics platform that helps address critical challenges in reputed company and productivity throughout the reputed company. It was founded in 2012, and is headquartered in San Francisco, California, USA, with a workforce of 51-200 employees. Its website is http://sightmachine.com.
  • Apply To This Job

    Similar Jobs

    [Remote] Data Engineer / Sr. Data Engineer (Big Query)

    Remote, USAFull-time

    [Remote] Senior Analyst, Procurement Data & Systems

    Remote, USAFull-time

    [Remote] Finance Counselor I

    Remote, USAFull-time

    [Remote] Director, Logistics Operations & Transformation

    Remote, USAFull-time

    [Remote] Senior Financial Analyst

    Remote, USAFull-time

    [Remote] Bilingual Spanish Customer Service Coordinator- Call Center

    Remote, USAFull-time

    [Remote] Health Solutions Legal Consultant (AVP) - Compliance & Policy Consulting Team

    Remote, USAFull-time

    [Remote] reputed company Backend Engineer

    Remote, USAFull-time

    [Remote] Senior DevOps Engineer - CI/CD & Kubernetes Platform

    Remote, USAFull-time

    [Remote] Finance Operations Coordinator

    Remote, USAFull-time

    reputed company Remote Customer Service Representative – reputed company Chat Support Team

    Remote, USAFull-time

    Junior Full Stack Data Entry Clerk – Remote Work Opportunity for Career reputed company and reputed company Development in Data Management Services at arenaflex

    Remote, USAFull-time

    Quantitative Analyst; Remote

    Remote, USAFull-time

    reputed company Full-Time RN Case Manager - Home Health Services in Huntsville, AL - Competitive Salary and Comprehensive Benefits

    Remote, USAFull-time

    Specialist, Clinical Informatics-Float

    Remote, USAFull-time

    Mental Health Therapist – Women's Health - 1099 Remote Contract

    Remote, USAFull-time

    [Remote] reputed company Manager

    Remote, USAFull-time

    RN Health Care Facility Surveyor - Rhode reputed company

    Remote, USAFull-time

    Product Manager, Payments Compliance

    Remote, USAFull-time

    [Remote] Sr. Core Backend Engineer

    Remote, USAFull-time