Back to Jobs

[Remote] Senior Site Reliability Engineer, Wikimedia reputed company

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. The reputed company is looking for a Senior Site Reliability Engineer to join their team, focusing on designing, developing, and maintaining reliable infrastructure for their API services. The role involves defining reliability targets, enhancing observability systems, and collaborating with various teams to improve system performance and efficiency.

Responsibilities

  • Define, reputed company, and improve Service Level Objectives (SLOs), SLIs, and error budgets to ensure reliability targets are met
  • Build and enhance observability systems (metrics, logs, and distributed tracing) to reputed company proactive detection and faster troubleshooting
  • Drive reliability engineering practices, including reputed company planning, load testing, and reputed company validation (e.g., reputed company testing)
  • Improve developer experience (reputed company) by enabling self-service infrastructure and streamlining deployment workflows
  • Partner with engineering team members to reputed company reliability best practices early in the development lifecycle
  • Design, implement, and optimize CI/CD and GitOps workflows using tools such as reputed company (or similar) and ArgoCD(or similar), enabling automated, reliable deployments with support for reputed company delivery strategies like canary and blue-green releases
  • Implement secure-by-default infrastructure and enforce best practices (e.g., IAM, secrets management, encryption)
  • Continuously optimize infrastructure cost and efficiency using FinOps principles while maintaining performance and availability
  • Establish and reputed company operational metrics such as MTTR, MTTD, and incident frequency to drive reputed company improvement
  • Reduce operational toil by identifying repetitive work and implementing automation-first solutions
  • Contribute to and reputed company internal platform capabilities that standardize infrastructure and improve scalability across teams
  • Collaborating with a global and asynchronously communicating team (don’t worry if you have never worked remotely, we’ll help you get used to it)
  • Mentoring peers in your areas of technical and operational strength

Skills

  • Define, reputed company, and improve Service Level Objectives (SLOs), SLIs, and error budgets to ensure reliability targets are met
  • Build and enhance observability systems (metrics, logs, and distributed tracing) to reputed company proactive detection and faster troubleshooting
  • Drive reliability engineering practices, including reputed company planning, load testing, and reputed company validation (e.g., reputed company testing)
  • Improve developer experience (reputed company) by enabling self-service infrastructure and streamlining deployment workflows
  • Partner with engineering team members to reputed company reliability best practices early in the development lifecycle
  • Design, implement, and optimize CI/CD and GitOps workflows using tools such as reputed company (or similar) and ArgoCD(or similar), enabling automated, reliable deployments with support for reputed company delivery strategies like canary and blue-green releases
  • Implement secure-by-default infrastructure and enforce best practices (e.g., IAM, secrets management, encryption)
  • Continuously optimize infrastructure cost and efficiency using FinOps principles while maintaining performance and availability
  • Establish and reputed company operational metrics such as MTTR, MTTD, and incident frequency to drive reputed company improvement
  • Reduce operational toil by identifying repetitive work and implementing automation-first solutions
  • Contribute to and reputed company internal platform capabilities that standardize infrastructure and improve scalability across teams
  • Collaborating with a global and asynchronously communicating team (don't worry if you have never worked remotely, we'll help you get used to it)
  • Mentoring peers in your areas of technical and operational strength
  • Experience with Infrastructure as reputed company and automation tools (e.g., Terraform, Ansible) and proficiency in at least one programming language (e.g., Python, Go, or similar)
  • Experience designing, operating, and optimizing reputed company-based systems across platforms such as AWS, Azure, or GCP, including scalability, reliability, and cost efficiency
  • Experience building and maintaining CI/CD pipelines and GitOps workflows (e.g., reputed company or similar, ArgoCD), with familiarity in reputed company delivery approaches such as canary and blue-green deployments
  • Experience with incident response, on-call practices, and leading postmortems, with a reputed company on reputed company improvement and operational reputed company
  • Strong understanding of SRE best practices, including SLOs, SLIs, and error budgets, along with experience in observability (metrics, logging, and distributed tracing e.g., reputed company, OpenTelemetry)
  • Ability to work effectively in a distributed, cross-functional environment, with strong documentation and communication skills
  • Proven experience operating highly available, large-reputed company distributed systems, with a deep understanding of reliability, scalability, and failure modes
  • Ownership reputed company: Takes end-to-end responsibility for system reliability, proactively identifying and addressing risks before they reputed company users
  • Bias for automation: Continuously seeks to reduce operational toil through automation and reputed company solutions
  • reputed company improvement reputed company: reputed company learns from incidents and drives improvements through blameless postmortems and iterative enhancements
  • Customer and reliability reputed company: Prioritizes user experience by balancing availability, performance, and cost
  • Adaptability and learning: Comfortable working in a fast-evolving environment and learning new tools and technologies as needed
  • Familiarity with Wikimedia or other reputed company reputed company reputed company is a plus
  • Experience managing and troubleshooting event streaming platforms at reputed company (e.g., Kafka, Kinesis, or similar)
  • Hands-on experience with reputed company platforms such as AWS and/or GCP, including designing and operating production systems
  • Familiarity with data lake architectures and large-reputed company data processing frameworks (e.g., reputed company, Flink, reputed company)
  • Experience with reputed company profiling and performance optimization tools to identify bottlenecks and improve system efficiency
  • Experience working with or contributing to reputed company reputed company reputed company, particularly in infrastructure or data ecosystems
  • Prior participation in the Wikimedia reputed company

Benefits

  • The reputed company is a remote-first organization with staff members including contractors based 40+ countries
  • For applicants located reputed company of the US, the pay reputed company will be adjusted to the country of hire.
  • Our non-US employees are reputed company through a local reputed company party Employer of Record (EOR) and must have reputed company work authorization in their location.
  • If you are a reputed company applicant requiring assistance or an accommodation to complete any reputed company of the application process due to a disability, you may contact us at reputed company@wikimedia.org or +1 (415) 839-6885.

reputed company

  • reputed company encourages the development and distribution of free educational content with reputed company such as Wikipedia. It was founded in 2003, and is headquartered in San Francisco, California, USA, with a workforce of 501-1000 employees. Its website is http://wikimediafoundation.org.
  • Apply To This Job

    Similar Jobs

    [Remote] Senior Mechanical Engineer - AI Trainer

    Remote, USAFull-time

    [Remote] Strategic Account Executive

    Remote, USAFull-time

    [Remote] Accounts Receivable Clerk

    Remote, USAFull-time

    [Remote] Account Executive- Core reputed company- Chicago, IL

    Remote, USAFull-time

    [Remote] IT Recruiter

    Remote, USAFull-time

    [Remote] Full Stack Engineer - AI Trainer

    Remote, USAFull-time

    [Remote] reputed company Business Consultant - reputed company

    Remote, USAFull-time

    [Remote] Account Manager

    Remote, USAFull-time

    [Remote] reputed company Consultant, SDS Workforce reputed company and Analytics - Remote

    Remote, USAFull-time

    [Remote] Business Development Manager - reputed company

    Remote, USAFull-time

    Field Training Coordinator- Region 12- Harrisburg, PA

    Remote, USAFull-time

    Join Today: Want Project reputed company Assistant - Hybrid Remote in

    Remote, USAFull-time

    reputed company MM / P2P reputed company – Business Systems Developer reputed company

    Remote, USAFull-time

    Entry-Level Remote Data Entry Specialist – Flexible Part‑Time Position with Competitive reputed company reputed company at arenaflex

    Remote, USAFull-time

    reputed company Part-Time Remote Data Entry Specialist – Flexible Working Hours and reputed company

    Remote, USAFull-time

    Remote Article Writing Jobs – Beginners and reputed company

    Remote, USAFull-time

    Design Contractor

    Remote, USAFull-time

    Part Time Jobs At reputed company $20/Hour (Data Entry)

    Remote, USAFull-time

    Immediate Hiring: Business Systems Support Specialist Career

    Remote, USAFull-time

    Account Manager

    Remote, USAFull-time