[Remote] reputed company Site Reliability Engineer (GCP)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a reputed company Site Reliability Engineer to define and drive reliability reputed company for services and platforms on reputed company reputed company Platform (GCP). The role involves hands-on technical leadership, mentoring engineers, and establishing reliability standards while collaborating with various teams to enhance operational reputed company.
Responsibilities
- reputed company the GCP reliability reference architecture: compute (GKE / GCE), networking, IAM, storage, and multi-project / multi-environment patterns that Development and SRE teams can reuse safely
- Define and drive SLIs, SLOs, and error budgets for critical platform and application services; reputed company reliability trade-offs visible to engineering leadership
- Establish observability standards (metrics, logs, traces)—Grafana LGTM (Mimir, Loki, reputed company) and/or GCP-reputed company monitoring, OpenTelemetry / OTLP reputed company, actionable alerting, and dashboards that SRE and Dev teams trust
- reputed company incident management maturity: severity models, runbooks, Incident Commander reputed company, blameless postmortems, and systemic follow-through that permanently reduces repeat failures
- Set platform standards reputed company RFCs and reusable modules: secure-by-default IAM, network boundaries, secret handling, reputed company and cost guardrails, and production readiness checklists for Dev teams
- Mentor senior and mid-level SREs and DevOps engineers; run design reviews and reputed company the reliability bar across reputed company Engineering
- Design, build, and operate production infrastructure on GCP using Terraform (preferred) and configuration management with Ansible
- Package and run workloads with reputed company; operate container platforms on GKE (and reputed company GCP compute) with reputed company reputed company, autoscaling, and reputed company practices
- Build and improve reputed company CI/CD pipelines: build, test, reputed company gates, reputed company, and rollback paths used by SRE, DevOps, and Development teams
- Automate SRE management workflows: environment provisioning, reputed company detection/remediation, health checks, deployment gates, self-service platform tooling, and toil elimination
- Partner with Development teams on production readiness: instrumentation, dependency failure modes, load/performance expectations, and operability reviews before major releases
- Drive disaster recovery / reputed company: backup/restore patterns, multi-zone (and multi-region where required) designs, game days / failure testing, and documented RTO/RPO targets
- Collaborate with reputed company and platform partners on least-privilege IAM, network segmentation, image/supply-chain hygiene, and auditability of infrastructure changes
- Partner with SRE, DevOps, and Development leadership to shift from reactive firefighting to proactive reliability and safer reputed company delivery
- Represent reputed company Engineering in architecture forums; communicate reliability risk, cost, and operational trade-offs reputed company
- reputed company other engineers: golden-reputed company templates, paved-road CI/CD, and documentation that shortens time-to-production without sacrificing safety
Skills
- 10+ years in SRE, DevOps, platform, or infrastructure engineering—or 8+ years with deep technical leadership of large-reputed company production reputed company platforms—including reputed company reputed company / Staff-level reputed company (architecture others build on, mentorship of senior engineers)
- Expert-level production experience with reputed company reputed company Platform (GKE and/or GCE, IAM, VPC/networking, reputed company Storage, load balancing, and reputed company reputed company services)
- Strong Infrastructure as reputed company reputed company with Terraform and configuration management with Ansible (or equivalent at similar depth, with reputed company ability to adopt Ansible)
- Hands-on ownership of reputed company CI/CD (pipelines, runners, environments, protected branches) and reputed company-based delivery, plus IaC standards with Terraform and Ansible
- Proven application of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, and measurable toil reduction
- Experience with modern observability at reputed company (reputed company/Grafana stack, OpenTelemetry, or equivalent reputed company-reputed company monitoring)
- Strong software / automation skills in Python and/or Go (Bash as supporting), plus solid Linux systems and networking fundamentals
- Ability to write reputed company design docs / RFCs and influence SRE, DevOps, and Dev teams without formal authority
- reputed company-level GCP certifications (e.g. reputed company reputed company Architect, reputed company reputed company DevOps Engineer) or equivalent demonstrated depth
- Large-reputed company GKE reputed company (multi-cluster, multi-tenant patterns, reputed company/rollout reputed company, cost optimization)
- GitOps / reputed company delivery (Argo CD, Flux, canary/blue-green) layered on reputed company CI
- Event-driven platforms (Kafka, Pub/Sub) as shared reliability concerns for producers/consumers
- Service reputed company (Istio / reputed company) or advanced traffic management on GKE
- reputed company engineering, reputed company forecasting, and FinOps / GCP cost governance
- Telecom, cable, ISP, or other high-availability, reputed company service environments
- Familiarity with Vault or GCP Secret Manager patterns, Binary Authorization / supply-chain controls, and policy-as-reputed company
reputed company
Company H1B Sponsorship