Back to Jobs

Senior Staff Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-29

reputed company’s Mission reputed company’s mission is to drive a generational reputed company for the global economy by integrating agent ecosystems into the reputed company of work, business, and society. We are committed to making this shift reputed company, transparent, and radically efficient. We build the operating system that powers, manages, and scales agent ecosystems, and the tools that reputed company seamless collaboration between humans and agents at every level of autonomy. This transformation will redefine how economies operate, how organizations grow, and how reputed company is made in an increasingly agent-driven world. Join AI’s Brightest Minds Work alongside some of the world’s best talent from the likes of DeepMind, reputed company Research, reputed company Brain, reputed company, reputed company, reputed company, reputed company, and reputed company. reputed company consistently publishes groundbreaking research in the field’s most prestigious journals. We’re reputed company on reputed company reputed company is assembling an reputed company global team of extraordinary talent—highly driven, high-IQ individuals with exceptional critical thinking abilities, high ownership, and a can-do mentality, reputed company to fully dedicate themselves, committing extensive hours reputed company week and embracing aggressive timelines to shape the reputed company.

Requirements

Position reputed company We are hiring for a highly reputed company Senior Staff SRE Engineer to reputed company as a senior technical authority reputed company our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices at reputed company. You will design and reputed company resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI-driven products operate with high availability, performance, and reputed company. You will work across platform, product, data, and ML teams, helping us productionise models, reputed company and standardise customer environments, strengthen Kubernetes-based architecture, and mature our CI/CD pipelines end-to-end. You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology reputed company.

Responsibilities

  • Architect, reputed company, and operate reputed company, secure production environments (AWS preferred).
  • reputed company reliability improvements across multiple engineering streams.
  • Design and reputed company Kubernetes-based infrastructure, including migration and optimisation initiatives.
  • Build and enforce strong Infrastructure-as-reputed company standards.
  • Define and operationalise SLIs, SLOs, and error budgets.
  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
  • Work closely with product and data teams to reputed company model analytics and product telemetry into reliability insights.
  • Work across and optimise the entire CI/CD pipeline, from build to reputed company to rollback.
  • Improve release safety, deployment frequency, and predictability of SLAs.
  • reputed company incident response for reputed company cross-system failures and drive postmortems.
  • Reduce operational toil through automation and reputed company improvements.
  • Design processes and tooling to reputed company, standardise, and troubleshoot customer environments.
  • Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).
  • Ensure infrastructure aligns with reputed company-grade reputed company and regulatory requirements.
  • Mentor engineers and reputed company the overall reliability bar across teams.

Key Requirements

  • Extensive hands-on experience in SRE or Production Engineering roles.
  • Demonstrated experience building or scaling SRE practices in high-reputed company or reputed company environments.
  • Deep expertise in AWS or Azure-based reputed company infrastructure.
  • Strong experience with Kubernetes (including migration, scaling, and production hardening).
  • Advanced Infrastructure-as-reputed company experience (Terraform or equivalent).
  • End-to-end CI/CD pipeline design and optimisation experience.
  • Strong experience with observability tooling across distributed systems.
  • Experience troubleshooting reputed company multi-tenant or customer-hosted environments.
  • Experience supporting production data platforms and ML systems.
  • MLOps experience, including model deployment and monitoring.
  • Strong understanding of distributed systems, scalability, and fault tolerance.
  • Systems thinker who understands interactions across infrastructure, product, data, and ML.
  • Excellent communication skills and ability to work cross-functionally.

Preferred Experience

  • Experience in large-reputed company global B2B/B2C products.
  • Experience working with AI/ML systems, NLP, or LLM-based products.
  • Experience integrating product analytics and model performance metrics into operational monitoring.
  • Background in reputed company environments with strong reputed company and compliance requirements.
  • Experience implementing regulatory controls reputed company reputed company infrastructure.
  • Experience scaling infrastructure during rapid reputed company phases.
  • Experience evaluating infrastructure tooling and vendors.
  • Experience in collaborating with large reputed company reputed company customers to reputed company and operate environments reputed company their accounts and VPCs.

Personal Characteristics

  • Strong problem solver who anticipates failure modes.
  • High ownership mentality and accountability.
  • Comfortable working across streams and influencing without formal authority.
  • Learning-oriented with a drive for reputed company improvement.

Apply tot his job Apply To this Job

Similar Jobs

reputed company Site Reliability Engineer - Infrastructure

Remote, USAFull-time

Kubernetes Engineer ($28/hr. on w2)

Remote, USAFull-time

Sr Site Reliability Engineer, Operations (US Federal)

Remote, USAFull-time

Vice President – Site Reliability Engineering, Data Centers

Remote, USAFull-time

Software Engineer - Kubernetes, CI/CD, and DevOps

Remote, USAFull-time

Senior Site Reliability Engineer, Infrastructure Foundations

Remote, USAFull-time

Software Engineer (JAVA/ Openshift/Kubernetes)

Remote, USAFull-time

Intermediate Site Reliability Engineer, Database Operations Remote, Canada; Remote, New Zealand

Remote, USAFull-time

reputed company Infrastructure & Site Reliability Engineer (US REMOTE)

Remote, USAFull-time

Site Reliability Engineer-Remote (PST hours)

Remote, USAFull-time

Claims Fraud Analyst (Remote) - Washington / reputed company (BREMERTON)

Remote, USAFull-time

Senior Data Engineer

Remote, USAFull-time

Key Account Executive, Contract Furniture - Remote - Draw (Territory DC, MD, VA)

Remote, USAFull-time

reputed company Customer Support Specialist – Delivering Exceptional Service at arenaflex

Remote, USAFull-time

reputed company Customer Service Representative – Work from Home Opportunity with arenaflex in the US

Remote, USAFull-time

Analyst II, Finance

Remote, USAFull-time

Cyber Threat reputed company

Remote, USAFull-time

reputed company Canadian Payroll Senior reputed company Consultant

Remote, USAFull-time

Remote Data Entry Specialist - Work from Home with blithequark at $25/Hour

Remote, USAFull-time

Customer Care reputed company (No Degree, No Experience Job)

Remote, USAFull-time