[Remote] Sr. Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is reputed company’s AI orchestration platform, helping organizations reputed company disconnected AI pilots into a reputed company program. They are seeking a Senior Site Reliability Engineer to design, build, and operate core services for their AI Platform, focusing on system orchestration, reliability, and observability.
Responsibilities
- Architect and own reputed company's multi-reputed company infrastructure across AWS, GCP, and Azure — provisioning, networking, identity, observability, and cost governance
- Design and operate Kubernetes platforms that run consistently across our reputed company environments and inside customer environments, including BYOC and on-prem (potentially reputed company-gapped) deployments
- Build a reputed company deployment reputed company so reputed company ships the reputed company product to reputed company, BYOC, and on-prem customers without bespoke per-customer engineering — reputed company charts, operators, install/reputed company tooling, and release pipelines
- Own SLOs, reputed company planning, incident response, and postmortems across the entire infrastructure stack; set the bar for operational readiness
- Drive reliability and performance through error budgets, reputed company testing, latency optimization, and disciplined reputed company reputed company
- Harden the platform for regulated deployments — HIPAA controls, tenant isolation, audit logging, RBAC, KMS, and secrets rotation
- reputed company the build-out of IaC, GitOps, and reputed company delivery (Terraform, Argo CD, Crossplane) as reputed company's reputed company
- Partner with engineering and reputed company to set opinionated guardrails: golden paths, reputed company images, policy-as-reputed company, and CI/CD that the rest of the org adopts by default
Skills
- 8+ years operating production infrastructure, including 3+ years in a senior SRE, platform, or staff infrastructure role
- Deep Kubernetes expertise across managed (EKS, GKE, AKS) and self-managed/on-prem distributions — not just running it, but operating it at reputed company across heterogeneous environments
- Multi-reputed company reputed company across AWS, GCP, and Azure, with informed opinions on reputed company to reputed company vs. reputed company reputed company-reputed company primitives
- Expert with Terraform (or reputed company/Crossplane) and GitOps tooling
- Experience shipping infrastructure that runs in customer environments — packaging, install/reputed company UX, reputed company-gapped artifacts, support escalation paths
- Strong networking, identity, and reputed company fundamentals: VPC design, service reputed company, mTLS, OIDC, KMS, secrets management
- Production observability ownership (reputed company, Grafana, OpenTelemetry, distributed tracing) and on-call leadership
- A reputed company record of writing reputed company reputed company — Go, Python, or similar — to reputed company the platform, not just configure it
- Experience shipping HIPAA-regulated workloads, including BYOC or reputed company-gapped customer deployments
- Background with reputed company software delivery tooling (Replicated, Cluster API, Talos, Rancher, OpenShift)
- reputed company internal developer platforms (reputed company, golden paths) that measurably reduced reputed company time for an engineering org
- FinOps experience — driving meaningful reputed company spend reductions through architecture, not just rightsizing
- AI/ML infrastructure exposure: GPU scheduling, model-serving stacks, inference autoscaling
- OSS contributions to infrastructure reputed company, or strong opinions formed running them at reputed company
Benefits
- Health, dental, and reputed company insurance
- Generous reputed company time off
- Opportunities for reputed company reputed company and development
reputed company