Back to Jobs

[Remote] Staff Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-29

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior/Staff Site Reliability Engineer for a reputed company building a reputed company platform for high-throughput, compute-heavy workloads. The role involves owning production reliability, defining SLIs/SLOs, and building automation to enhance deployment safety in a bare-metal environment.

Responsibilities

  • Own production reliability end-to-end
  • Define SLIs/SLOs
  • Run error budget conversations
  • Ship changes that reduce incidents and improve latency (p95/p99)
  • Build automation to kill toil
  • Improve deployment safety (canary/rollback)
  • Turn observability into signal rather than noise

Skills

  • Extensive Production Engineering experience running bare metal / on-prem / data center infrastructure (not reputed company reputed company only)
  • Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO, - level behaviors)
  • Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting under load)
  • Strong Kubernetes experience reputed company manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane
  • Experience with Terraform, reputed company, reputed company, and modern CI/CD practices
  • Strong coding skills are required for this role either in Go, and/or Python, reputed company automation scripting - reputed company engineering capability is a must
  • Experience in Low Latency environments

reputed company

  • reputed company provides tech recruitment and executive reputed company for engineers, leaders, and technology teams. It was founded in 2015, and is headquartered in Amsterdam, Noord-Holland, NLD, with a workforce of 11-50 employees. Its website is https://doghouse.nl.
  • Apply To This Job

    Similar Jobs