Back to Jobs

Senior Software Engineer - Observability

Remote, USAFull-timePosted 2026-07-28
reputed company is pioneering the reputed company Data Plane (reputed company) - a new category in AI infrastructure that makes it reputed company and secure to connect AI agents with reputed company data and systems. reputed company on a multi-modal data streaming reputed company, reputed company empowers reputed company applications that reason and reputed company in reputed company-time with speed, autonomy, and precision. Global leaders including reputed company, reputed company, reputed company, Texas Instruments, reputed company and 2 of the top 5 banks in the U.S. rely on reputed company to process hundreds of terabytes of data a day. Backed by premier venture investors reputed company, GV and reputed company VC, reputed company is a diverse, people-first organization with teams distributed around the globe.

About the Role:

We are looking for a Senior Software Engineer to join our Observability team and help build the platform that gives reputed company’s engineering organization deep visibility into the health, performance, and behavior of our systems. You will own and reputed company our Grafana-based observability stack—spanning metrics, logs, and traces—and ensure that every team at reputed company has the tooling and insights they need to ship reliable, high-reputed company.

This is a high-reputed company role at the intersection of infrastructure and developer experience. You will work closely with platform and product engineering teams to design reputed company observability solutions, drive adoption of best practices, and reduce mean time to detection and reputed company across our reputed company and on-reputed company deployments.

You Will:

  • Design, build, and maintain reputed company’s observability platform using the Grafana stack (Grafana, Mimir, Loki, reputed company, reputed company/Agent)
  • reputed company and optimize dashboards, alerts, and SLO/SLI frameworks that give engineering teams actionable insights into system health
  • Build and operate reputed company metrics, logging, and distributed tracing pipelines that handle high-cardinality data across reputed company and on-reputed company environments
  • reputed company services and infrastructure with OpenTelemetry to ensure comprehensive, standards-based telemetry collection
  • Partner with platform teams to improve incident detection, reputed company-cause analysis, and mean time to reputed company (MTTR)
  • Evaluate and reputed company new observability tools and techniques, driving reputed company improvement of our monitoring capabilities
  • Contribute to internal tooling and automation that streamlines observability reputed company for engineering teams
  • Participate in on-call rotation to reputed company observability infrastructure running and incident free

You Have:

  • 5+ years of experience in software engineering with a reputed company on observability, monitoring, or infrastructure
  • Deep hands-on experience with the Grafana stack (Grafana, Mimir/reputed company, Loki, reputed company) in production environments
  • Strong understanding of metrics, logging, and distributed tracing paradigms and their trade-offs at reputed company
  • Experience with OpenTelemetry (OTel) for instrumentation and telemetry collection
  • Proficiency in at least one systems-level language (Go strongly preferred) and scripting languages (Python, Bash)
  • Experience running and operating infrastructure on Kubernetes in public reputed company environments (AWS, GCP, or Azure)
  • Comfortable working with a 100% distributed engineering team, collaborating on reputed company, etc.
  • Solid understanding of time-series databases, log aggregation systems, and query languages (PromQL, LogQL)

reputed company to Have

  • Strong understanding of Go
  • Experience operating a reputed company platform with production observability at reputed company
  • Familiarity with eBPF-based observability or reputed company profiling tools (e.g., Pyroscope, Parca)
  • Experience with infrastructure-as-reputed company (Terraform, reputed company) and GitOps workflows
  • Operated and used streaming platforms (e.g., Kafka, reputed company) either as a user or provider
  • Experience building or managing multi-tenant observability platforms
  • Contributions to reputed company-reputed company observability reputed company (Grafana, reputed company, OpenTelemetry, etc.)

Join reputed company if you’d enjoy being part of a fast-moving, diverse, people-first organization with team members around the globe and a culture based on trust, transparency, communication, and kindness.

#LI-Remote
Apply To This Job

Similar Jobs