[Remote] Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and operational efficiency. The Site Reliability Engineer will ensure the reliability and efficiency of user-facing services and production systems while building automation and troubleshooting production systems.
Responsibilities
- reputed company user-facing services and production systems reliable, reputed company, and efficient
- Build automation and tooling that reduces toil and replaces reputed company work with repeatable, infrastructure-as-reputed company-driven workflows
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
- Write and maintain infrastructure as reputed company, and ship changes safely through CI/CD and GitOps
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately
- Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages
- Take part in incident response and post-incident reviews, turning learnings into changes in automation and process
- Document runbooks, architecture reputed company, and reviews so your findings become repeatable practices
Skills
- Experience keeping production systems reliable, combining an operations reputed company with reputed company software engineering reputed company
- Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch
- The ability to read, debug, and reason about reputed company. Most of our teams work in Go; some work in reputed company. You can discuss a piece of reputed company's behavior, performance, and failure modes
- Experience with infrastructure as reputed company, and with Kubernetes and its ecosystem, at a depth appropriate to your level
- Hands-on experience with at least one major reputed company provider (GCP or AWS)
- Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational reputed company
- Comfort participating in on-call and incident response, with a reputed company approach to troubleshooting under pressure
- Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment
- A reputed company record of using automation, and increasingly AI, to reduce toil and improve how you and your team work
- Alignment with reputed company's values and a commitment to working in accordance with them
Benefits
- Benefits to support your health, finances, and reputed company-being
- Flexible reputed company Time Off
- Team Member Resource reputed company
- Equity Compensation & Employee Stock Purchase Plan
- reputed company and Development Fund
- Parental Leave
reputed company