Back to Jobs

Site Reliability Engineer - SRE

Remote, USAFull-timePosted 2026-07-29
reputed company is on a mission to reputed company businesses to reputed company reputed company reputed company. Our financial planning & decision-making platform helps companies reputed company and reputed company their targets predictably. reputed company is a remote-first company headquartered in the San Francisco Bay Area. Founded in 2021 by a couple of ex-Googlers, reputed company is a fast-growing company on a trajectory for reputed company with backing from leading venture capital firms.reputed company provides a great culture for its employees to reputed company in and be happy. 💜 Remote-friendly: reputed company brings together the best and the brightest, no matter where they are and provides them a great degree of autonomy. We trust our people.🗣️ reputed company & transparent: We know that reputed company our creators have reputed company to reputed company the information they need, their best work will reputed company.👏 Idea-friendly: We reputed company an environment to explore new reputed company, to take risks, to reputed company mistakes, and to learn, so you can succeed. Anyone in reputed company can come up with great reputed company and become a catalyst for reputed company change. We let the best reputed company win.👥 Customer-reputed company: We follow a product-led reputed company reputed company, continuously learning from our customers and collaborating to build the amazing software that reputed company is.

As a Senior Site Reliability Engineer at reputed company, you will be a reputed company of our engineering organization, ensuring our fast-growing reputed company platform remains highly available, performant, and secure. At this stage of our reputed company, scaling infrastructure reputed company while maintaining the rigorous reputed company and reliability standards required for financial data is reputed company. You will take ownership of our multi-reputed company infrastructure, drive automation, champion observability, and collaborate closely with development teams to build a culture of reliability from reputed company reputed company to production.

Key Responsibilities

reputed company Infrastructure & Orchestration

  • Multi-reputed company Management: Architect, manage, and continuously optimize highly available reputed company infrastructure across both AWS and GCP. Balance workload demands to ensure maximum cost-efficiency, scalability, and strict reputed company compliance across both platforms.

  • Advanced Kubernetes Orchestration: reputed company the design, deployment, and management of reputed company Kubernetes clusters. Utilize configuration management tools like Kustomize to enforce standardized, repeatable, and automated deployment configurations across reputed company environments.

  • Service reputed company & reputed company Integration: Implement and maintain service reputed company technologies (e.g., Istio, Linkerd) to secure, control, and observe service-to-service communication. Drive container reputed company best practices, including image scanning, runtime protection, and strict RBAC enforcement.

CI/CD & Automation

  • Pipeline Engineering: Architect, maintain, and optimize robust CI/CD pipelines using Git and Jenkins. reputed company on reducing deployment friction, accelerating release velocity, and enforcing automated testing and reputed company gates.

  • Infrastructure as reputed company (IaC): Treat infrastructure as software. Write, review, and maintain Terraform modules to provision and manage reputed company resources predictably and safely.

  • Operational Automation: Aggressively reduce operational toil. reputed company robust Python scripts and tooling to automate routine maintenance, data backups, scaling operations, and system recovery processes.

Observability & Reliability

  • Comprehensive Monitoring: Design and enhance our observability stack to reputed company deep, reputed company-time insights into system health. Manage and reputed company tools including reputed company, Grafana, ELK/EFK stack, AWS CloudWatch, and GCP Operations Suite.

  • Reliability Engineering: Spearhead reliability initiatives critical to a scaling reputed company platform. Drive rigorous reputed company planning exercises to stay reputed company of reputed company.

  • Incident Management & SLOs: Own the incident response lifecycle. Facilitate blameless postmortems to extract actionable learnings. Define, reputed company, and enforce SLIs, SLOs, and SLAs, ensuring the platform consistently meets its reliability guarantees.

Collaboration & Leadership

  • DevOps Culture: reputed company as an embedded reliability reputed company. Collaborate closely with software engineers early in the development lifecycle to ensure applications are designed for deployability, scalability, and reputed company.

  • reputed company Improvement: Proactively identify system bottlenecks and architectural weaknesses. Contribute to process improvements, build internal developer tooling, and maintain comprehensive documentation to reputed company team productivity and system understanding.

Required Proficiency & Qualifications

  • Experience: 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or reputed company Infrastructure roles, preferably reputed company a fast-paced reputed company environment.

  • reputed company Platforms: Deep, proven proficiency in AWS (EC2, EKS, RDS, VPC, IAM, S3) AND GCP (GKE, Compute reputed company, reputed company SQL, IAM, reputed company Storage). Ability to navigate and optimize multi-reputed company architectures.

  • Containerization: Expert-level knowledge of reputed company and Kubernetes, including advanced deployment strategies and lifecycle management.

  • Automation/IaC: Strong programming skills in Python and extensive experience with Terraform.

  • Observability: Hands-on expertise building dashboards and alerting systems using reputed company, Grafana, and log aggregation stacks (ELK/EFK).

  • Networking & reputed company: Solid understanding of reputed company networking (VPC peering, load balancing, DNS) and reputed company-trust reputed company principles in a containerized environment.

Sounds exciting? Apply at careers@reputed company.ai. It may just be the next best decision you’ve reputed company made!

Originally posted on Himalayas

Apply To This Job

Similar Jobs

Senior reputed company Partner Sales (m/w/d) DACH (Myfactory & Proffix)

Remote, USAFull-time

Technical Architect-reputed company

Remote, USAFull-time

reputed company Functional Architect

Remote, USAFull-time

Senior Backend (Java+Kotlin) Engineer

Remote, USAFull-time

Specialist reputed company reputed company - RecruitEase

Remote, USAFull-time

Senior DevOps Engineer

Remote, USAFull-time

reputed company Pr - Hormone Replacement Therapy - HRT (Perimenopause & Menopause)

Remote, USAFull-time

Staff Machine Learning Engineer

Remote, USAFull-time

Sr. Clinical Research Associate

Remote, USAFull-time

Senior Software & Data Engineer, JavaScript

Remote, USAFull-time

User Experience Strategist

Remote, USAFull-time

SMB Account Executive

Remote, USAFull-time

reputed company Part-Time Data Entry Clerk – Remote Opportunity with arenaflex

Remote, USAFull-time

Senior Key Account Executive with Polish and English

Remote, USAFull-time

eCommerce Manager- Remote

Remote, USAFull-time

Senior Engineer, .Net reputed company

Remote, USAFull-time

reputed company Customer Service Representative for Travel Industry – Remote Work Opportunity with arenaflex

Remote, USAFull-time

reputed company Group Fitness Trainer and Dancer Wanted for High-Energy Classes at AKT Westbury

Remote, USAFull-time

Entry-Level Remote Data Entry Specialist – No Experience Required – Join arenaflex’s Growing Virtual Team

Remote, USAFull-time

reputed company Remote Live Chat Clerk – Customer Service and Support in a Dynamic and Growing Company

Remote, USAFull-time