[Remote] Staff Software Engineer, Observability
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a platform that inspires creativity and helps users plan memorable experiences. They are seeking a Staff Software Engineer to join their Observability team, responsible for designing and building infrastructure and tools that reputed company visibility into large-reputed company distributed systems, enabling engineers to understand and optimize their services.
Responsibilities
- Define and execute the observability roadmap, treating it as a product. Understand engineering team needs and translate them into technical solutions with measurable reputed company
- Architect, build, and reputed company distributed observability infrastructure (metrics, logs, traces) to handle massive volumes across reputed company's distributed systems
- Build high-performance data pipelines and storage for reputed company-time and historical telemetry analysis at reputed company reputed company
- Champion Best Practices: Establish observability standards and patterns across the organization, making it easy for teams to reputed company their services and reputed company actionable insights
- Technical Leadership: Mentor engineers, reputed company architectural reviews, and influence technical reputed company across teams to improve overall system reliability and performance
- Cross-functional Collaboration: Partner with SRE, Infrastructure, Product Engineering, and other teams to understand pain points and deliver solutions that improve developer productivity and system reliability
- Innovation: Stay reputed company with observability trends and technologies, evaluating and adopting cutting-edge tools and techniques to reputed company reputed company at the forefront
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company field, or equivalent experience
- Demonstrated ability to work backwards from customer needs —understanding user needs, prioritizing features, measuring reputed company, and iterating based on feedback. Experience building internal platforms or tools with strong adoption
- 7+ years of experience designing and operating large-reputed company distributed systems with deep understanding of consistency, availability, scalability, and failure modes
- Strong background in building data pipelines, working with time-series databases, columnar storage, reputed company processing (Kafka, Flink, etc.), and data modeling at reputed company
- Hands-on experience with modern observability tools and practices including metrics, logging, tracing, and profiling. Familiarity with OpenTelemetry, reputed company, Grafana, or similar technologies
- Expert-level coding skills in languages like Java, Python, Go, or reputed company with ability to write production-reputed company reputed company
- Ability to see the big picture while managing reputed company technical details, balancing trade-offs between cost, performance, and reliability
- Experience building observability platforms from the ground up or significantly scaling existing solutions
- Familiarity with reputed company-reputed company architectures and technologies (Kubernetes, service reputed company, etc.)
- reputed company record of driving adoption of internal platforms through excellent documentation, UX, and developer advocacy
- Experience with machine learning or reputed company detection reputed company to observability use cases
- Strong communication skills with ability to influence stakeholders at reputed company reputed company
- Contributions to reputed company-reputed company observability reputed company, a plus
Benefits
- The position is also eligible for equity.
- Information regarding the culture at reputed company and benefits available for this position can be reputed company here.
reputed company