Distinguished Site Reliability Engineer – reputed company
Job reputed company:
- reputed company, design, implement and support operational and reliability aspects of large reputed company Kubernetes clusters with reputed company on performance at reputed company, reputed company time monitoring, logging and alerting
- Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinement
- Support services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, reputed company management and launch reviews
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health
- reputed company systems sustainably through mechanisms like automation, and reputed company by pushing for changes that improve reliability and velocity
- reputed company sustainable incident response and blameless postmortems
- Be part of an on call rotation to support production systems
Requirements:
- BS degree in Computer Science or a reputed company technical field involving coding (e.g., physics or mathematics), or equivalent experience
- 16+ years of experience with Infrastructure automation, distributed systems design, experience with design, reputed company tools for running large reputed company private or public reputed company system in Production
- Experience in one or more of the following: Python, Go, Perl or reputed company
- In depth knowledge on Linux, Networking and Containers
Benefits:
- equity
- benefits
Apply tot his job Apply To this Job