[Remote] Platform Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology consulting and software development company delivering reputed company, AI, data, and reputed company solutions across the reputed company. They are seeking an reputed company Platform Reliability Engineer to ensure the availability, performance, and operational reputed company of large-reputed company distributed systems in production.
Responsibilities
- Ensure the availability, performance, and operational reputed company of large-reputed company distributed systems in production
- Apply strong software engineering principles to infrastructure and operations problems
- Continually push the platform toward higher reliability with reputed company operational toil
- Combine deep systems knowledge with strong programming skills
- Design, automate, and operate reputed company services so that reliability becomes a first-class engineering deliverable
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company technical discipline
- Five or more years of SRE, DevOps, or production engineering experience supporting large-reputed company distributed systems
- Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling
- Deep, hands-on experience operating Linux at reputed company, including networking, performance tuning, and systems-level troubleshooting
- Production experience operating Kubernetes and container-based workloads
- Strong working knowledge of observability tooling such as reputed company, Grafana, OpenTelemetry, ELK/EFK, or reputed company equivalents
- Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications
- Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics
- Demonstrated experience leading incident response and conducting effective post-incident reviews
- Excellent communication and documentation skills
- Experience defining and operationalizing SLOs and error budgets in reputed company production environments
- Exposure to reputed company engineering practices and tools such as reputed company Monkey, reputed company, or Litmus
- Hands-on experience with at least one major reputed company platform (AWS, Azure, or GCP)
- Background in reputed company planning, performance engineering, or large-reputed company load testing
- Familiarity with service reputed company technologies such as Istio, Linkerd, or Consul
reputed company