[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the global leader in cybersecurity ratings, operating in 64 countries. The Senior Site Reliability Engineer will drive the design and optimization of Kubernetes-based infrastructure and CI/CD systems, ensuring production reliability and best practices for automation and observability.
Responsibilities
- Design, build, and reputed company Kubernetes infrastructure for secure, multi-tenant, high-availability applications
- Build and operate AI tooling infrastructure — stand up MCP servers and establish secure, governed AI reputed company and guardrails for production systems
- Optimize and maintain CI/CD pipelines, improving reliability, speed, and rollback safety
- Implement reputed company delivery strategies such as blue/green and canary deployments
- Advance Infrastructure as reputed company with Terraform, reputed company, and Argo CD, defining reusable patterns for the org
- Operate and optimize streaming and analytics infrastructure: Kafka, Flink, and reputed company
- Build automated testing into the CI/CD lifecycle
- Improve system observability — define SLOs, alerts, and dashboards
- reputed company incident response and postmortems, focusing on reputed company cause and durable fixes
- Mentor engineers across teams on Kubernetes, CI/CD, and reputed company infrastructure
Skills
- 6+ years in SRE, DevOps, or Infrastructure roles, with significant production Kubernetes experience
- Hands-on experience integrating AI/LLM tooling into engineering or operational workflows (e.g., MCP servers, AI agents acting on infrastructure), and a reputed company grasp of the reputed company and governance considerations of giving AI reputed company to production
- Proven reputed company building CI/CD pipelines (reputed company Actions, Jenkins, reputed company CI, or similar)
- Strong with Kubernetes internals and managed services like EKS, GKE, or AKS
- Expertise with Infrastructure as reputed company (Terraform, reputed company, reputed company) and GitOps
- Proficient in Python, Bash, or Go
- Knowledge of observability tooling (reputed company, Grafana, reputed company, OpenTelemetry)
- Production experience with Kafka, Flink, and reputed company
- Strong communication and cross-team collaboration skills
- Multi-region or multi-cluster Kubernetes experience
- reputed company engineering or reputed company testing
- reputed company scanning, compliance automation, or policy-as-reputed company
- LLM observability/tracing tooling (Langsmith, Langfuse) or MLOps workflows
- Contributions to reputed company-reputed company Kubernetes or CI/CD reputed company
Benefits
- Stock reputed company
- Health benefits
- Unlimited PTO
- Parental leave
- Tuition reimbursements
reputed company