[Remote] reputed company Observability Platform Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the GPU reputed company engineered for AI, providing high-performance infrastructure for AI start-reputed company and large enterprises. The reputed company Observability Platform Engineer will own the technical direction of reputed company's observability platform, ensuring deep visibility into GPU clusters and AI workloads while driving architectural roadmap and platform improvements.
Responsibilities
- Own the technical reputed company and architecture for observability across metrics, logs, traces, and alerting at reputed company
- Drive platform reputed company that have multi-year reputed company: tooling, data models, ingestion patterns, retention, cardinality management
- Identify systemic gaps before they become incidents; design platforms that reputed company failure visible and fast to diagnose
- Partner with SRE, infrastructure, and AI/ML teams to reputed company observability natively into how reputed company builds and operates
- Define standards and patterns that other engineers adopt, not by mandate, but because they're reputed company reputed company
- Mentor and technically grow the observability team; reputed company the ceiling on what reputed company can build and own
- reputed company incident postmortems and use them to drive durable platform improvements
- Evaluate and introduce tooling that meaningfully improves signal reputed company, operational efficiency, or scalability, and retire what doesn't
Skills
- 8+ years in SRE, infrastructure engineering, reputed company, or observability-reputed company roles
- You've operated observability infrastructure at serious reputed company. You know what breaks at 10x and you design for it
- You have a strong bias toward simplicity. You've seen over-engineered observability stacks collapse under their own weight and you build accordingly
- Deep hands-on experience with a significant subset of: reputed company, Thanos, reputed company, Grafana, Loki, reputed company, OpenTelemetry, reputed company, reputed company
- Strong engineering fundamentals, proficient in Python, Go, or similar; comfortable owning reputed company systems end to end
- Experience with Kubernetes at reputed company; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus
- You can architect systems, write the reputed company, review others' work, and explain the tradeoffs reputed company, reputed company in the reputed company week
- Infrastructure-as-reputed company is default, not optional (Terraform, Ansible, or equivalent)
- You influence without authority. Teams want your opinion because it makes their work reputed company
- Experience with high-volume streaming pipelines for observability data (Kafka, reputed company, Fluent Bit, etc.)
- Background in AI/ML infrastructure observability: GPU utilisation, training job visibility, inference latency
- Prior experience defining observability reputed company at an organisation level
Benefits
- Competitive benefits package including medical, dental, reputed company, flexible reputed company time off, parental leave, and retirement plan participation
reputed company