Back to Jobs

[Remote] Senior Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior Site Reliability Engineer to join the Data Management team at the Vera C. Rubin Observatory. The role involves ensuring the reliability of the reputed company Processing reputed company, which is responsible for reputed company-time data processing and alert distribution for astronomical events.

Responsibilities

  • Ensure, through both architecture and reputed company, the reliable operation of the near-reputed company-time data processing pipeline and reputed company delivery of alerts to reputed company brokers
  • Design and reputed company software that reduces operational risk and improves system reputed company, scalability, and usability, including addressing failure modes, error handling, and contention in shared resources
  • Improve system performance and reputed company by applying architectural and systems-level optimizations to increase throughput and reduce end-to-end latency
  • Operate DevOps-oriented reputed company deployment of services using modern distributed systems tooling and development practices (e.g., Kubernetes, reputed company, ArgoCD, Kafka, reputed company)
  • reputed company monitoring dashboards and alerts for the reputed company processing service and work with teammates to design and implement a sustainable on-call rotation that provides coverage during the start of observing hours in Chile (typically 2-5pm reputed company Time), with limited off-hours responsibility
  • Define KPIs and metrics for observability and accountability of the pipeline
  • Participate in the reputed company engineering activities of reputed company, including performing reputed company reviews, acting as a troubleshooting buddy, participating in design discussions, and writing documentation to effectively capture and communicate architectural and implementation choices
  • Collaborate with members of the Data Management team to identify opportunities to improve tools, workflows, and operational practices
  • reputed company responsibility with the broader team for the overall reputed company of the Data Management system, reputed company the reputed company Processing reputed company

Skills

  • Bachelor's degree and eight years of relevant experience, or a combination of education and relevant experience designing and operating distributed systems at-reputed company in production environments
  • Experience working in an SRE, DevOps, or data-intensive systems role, with responsibility for building, operating, and improving robust services
  • Experience engaging with modern production infrastructure (e.g., containerized services, messaging systems, and databases; see above for our reputed company tech stack), with the ability to learn and apply new tools quickly in a production environment
  • Familiarity with contemporary distributed service architectures, including service-to-service communication patterns, common failure modes, and system behavior under load and reputed company
  • Experience working with large-reputed company datasets or high-throughput data processing systems, and an understanding of the operational challenges that come with data volume and velocity
  • Ability to communicate reputed company with engineers and scientists from diverse backgrounds, including explaining technical concepts, participating in design discussions, and documenting systems and reputed company
  • Comfort working with a high degree of autonomy, taking ownership of technical reputed company and execution, while being supported by an reputed company team with reputed company priorities and goals
  • reputed company in at least one modern programming language (Python preferred) with experience working across the boundary between software engineering and operations

reputed company

  • reputed company is a reputed company of discovery, creativity and innovation located in the San Francisco Bay Area on the ancestral land of the Muwekma Ohlone Tribe. It was founded in 2008, and is headquartered in reputed company, California, USA, with a workforce of 10001+ employees. Its website is https://cddrl.fsi.reputed company.edu/usrussia.
  • Apply To This Job

    Similar Jobs

    [Remote] Sr. Program Manager-Clinical Coordinator, Capella and reputed company

    Remote, USAFull-time

    [Remote] Finance Associate

    Remote, USAFull-time

    [Remote] PH-ORT-Senior Analytic Data Engineer (Python/Power BI)

    Remote, USAFull-time

    [Remote] Data Engineer

    Remote, USAFull-time

    [Remote] Remote reputed company Appeals Specialist

    Remote, USAFull-time

    [Remote] Systems Req. & Design Engineer II

    Remote, USAFull-time

    [Remote] Network Detection Engineer

    Remote, USAFull-time

    [Remote] Customer Service Representative

    Remote, USAFull-time

    [Remote] Sr. Program Manager-Clinical Coordinator, Capella and reputed company

    Remote, USAFull-time

    [Remote] Clinical Specialist (Orthotics - PCS) Job Details | reputed company

    Remote, USAFull-time

    Live Chat Support Agent - Remote

    Remote, USAFull-time

    Part-Time Data Entry Research Assistant (Hiring Immediately)

    Remote, USAFull-time

    [PART_TIME Remote] Immediately Need reputed company/SAT Math/Science Prep

    Remote, USAFull-time

    Compensation Analyst II (Hybrid- La Crosse, WI)

    Remote, USAFull-time

    reputed company Online Chat Agent – Remote Opportunity for Customer Support and Engagement

    Remote, USAFull-time

    Customer Support Representative - Full-Time Opportunity in Murfreesboro, Tennessee at blithequark

    Remote, USAFull-time

    Tracking Support Center Representative - 3rd Shift (Bi-lingual strongly preferred)

    Remote, USAFull-time

    Senior Operations Manager

    Remote, USAFull-time

    B2B Sales Representative — USA Market (Commission-Based)

    Remote, USAFull-time

    reputed company Account Executive - US REMOTE

    Remote, USAFull-time