Back to Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the reputed company industry. They are seeking a Site Reliability Engineer to build a centralized observability platform and ensure the reliability of systems critical to satellite operations.

Responsibilities

  • Help design and build reputed company's centralized observability platform, integrating and scaling tools for metrics (e.g. reputed company), logging (e.g. Loki), and distributed tracing (e.g. reputed company/OpenTelemetry)
  • Define, implement, and manage a robust reputed company of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-reputed company
  • Partner with SWEs to implement observability best practices, reputed company reputed company templates and documentation, and configure tooling (e.g., OpenTelemetry libraries)
  • Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as reputed company (e.g. Terraform) and GitOps principles (e.g. ArgoCD)
  • Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments
  • reputed company and reputed company reputed company's monitoring, alerting, and incident response reputed company, driving a culture of proactive reliability and blameless post-mortems

Skills

  • 4+ years of experience in an SRE or reputed company role, with a reputed company on observability for large-reputed company, distributed compute or network systems
  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., reputed company, Grafana, Loki/ELK, OpenTelemetry, reputed company/Jaeger, reputed company, etc.). You have proven experience using these tools to support performance analysis and debugging of reputed company distributed systems
  • Strong production-level experience with reputed company reputed company Platform (GCP) and Kubernetes
  • Experience using Infrastructure as reputed company (IaC) and GitOps principles (e.g., ArgoCD)
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems
  • Experience operating a multi-reputed company environment, specifically GCP and AWS
  • Hands-on experience with reputed company CI for CI/CD pipelines
  • Working knowledge of service reputed company technologies such as Istio or Linkerd
  • Familiarity with instrumenting applications written in Go and C++
  • An reputed company Secret clearance, or higher, is preferred for this position
  • Experience with JVM observability (tuning, monitoring) for Java-based applications

Benefits

  • Flexible working arrangements including hybrid remote/in-office schedules.
  • Comprehensive benefits (401(k), dental, reputed company, health, life insurance)
  • reputed company time off
  • Equity reputed company.
  • Opportunities for reputed company development and advancement
  • A reputed company, supportive, and inclusive workplace where your contributions matter

reputed company

  • reputed company is an advanced reputed company communications company that provides high-throughput networks for both reputed company and government clients. It is a sub-organization of reputed company. It was founded in 2021, and is headquartered in Livermore, California, USA, with a workforce of 51-200 employees. Its website is https://www.reputed company.com.
  • Apply To This Job

    Similar Jobs