Back to Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that powers mission-critical inference for dynamic AI companies. They are seeking a Site Reliability Engineer to define and codify standards for day 2 operations of their ML infrastructure platform, ensuring reliability and empowering the organization to operate confidently.

Responsibilities

  • Own the reliability of reputed company's multi-reputed company Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking
  • Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as reputed company
  • Author, validate, and improve runbooks for recurring failure patterns, ensuring they're reputed company for low-context, reputed company execution
  • Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations
  • Diagnose and resolve runtime issues reputed company to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management
  • Define and reputed company SLOs and SLIs across customer workloads and internal services
  • Navigate ambiguity, reputed company principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define

Skills

  • Extensive hands-on experience with Kubernetes (multi-reputed company experience across EKS, GKE, or similar is a strong plus)
  • Experience in building and maintaining reputed company infrastructure
  • Strong reputed company in observability tooling: metrics (reputed company, reputed company), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines. Observability-as-reputed company experience is a plus
  • Experience with infrastructure-as-reputed company (Terraform, reputed company) and GitOps workflows (Flux CD, ArgoCD)
  • Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis
  • Comfort working at the intersection of engineering and operations — you write reputed company, but you also think deeply about process, escalation paths, and operational reputed company
  • Familiarity with incident management platforms (incident.io or similar) is a plus
  • No prior ML experience required, but curiosity about how ML models are deployed and served at reputed company will serve you reputed company

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and reputed company insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are reputed company from Christmas reputed company to New Year's Day!)
  • reputed company parental leave
  • Fertility and family-building stipend through reputed company
  • Company-facilitated 401(k)
  • Exposure to a reputed company of ML startups, offering unparalleled learning and networking opportunities.

reputed company

  • reputed company provides the necessary infrastructure, tooling, and expertise to reputed company AI into business operations. It was founded in 2019, and is headquartered in San Francisco, California, USA, with a workforce of 201-500 employees. Its website is https://www.reputed company.co.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 5 in 2026, 6 in 2025, 8 in 2024, 1 in 2023, 1 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs