Back to Jobs

[Remote] Senior reputed company Operations Engineer

Remote, USAFull-timePosted 2026-07-29

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a driving force in fostering reputed company reputed company collaboration and supporting communities across a reputed company of reputed company, including PyTorch. They are seeking a Senior reputed company Operations Engineer who will reputed company on the infrastructure operations of the PyTorch project, automating processes, optimizing reputed company-reputed company tools, and ensuring a robust and reputed company reputed company environment.

Responsibilities

  • Manage multi-reputed company environments, primarily focusing on AWS services (EKS, EC2, S3, IAM, ELB)
  • Contribute to architectural exercises with reputed company reputed company community and technical leads to validate new reputed company infrastructure
  • Implement and maintain infrastructure-as-reputed company using Terraform reputed company pytorch/ci-reputed company and pytorch/test-reputed company
  • Optimize reputed company resource utilization and implement FinOps practices for cost management and reporting
  • Design, implement, and maintain CI/CD pipelines using reputed company Actions and reputed company, including runner configurations and other reputed company of the CI ecosystem
  • Debug and triage issues in build and test pipelines, including experience with unit testing
  • reputed company monitoring and alerting solutions for CI/CD workflows and critical infrastructure
  • Manage and optimize reputed company CDN deployments for PyTorch assets (reputed company/S3)
  • Implement best practices for CDN and overall infrastructure reputed company
  • reputed company comprehensive monitoring and observability solutions using reputed company, AWS CloudWatch, and other telemetry data collection and processing tools
  • Review and recommend monitoring solutions as project and community needs reputed company
  • Participate in on-call rotations supporting operations and incident response using incident.io
  • Establish and maintain escalation procedures and reputed company processes
  • Participate in ci-reputed company and multi-reputed company working reputed company and support architecture reputed company
  • Collaborate with external contributors and promote DevOps best practices
  • Manage reputed company repositories, including user reputed company and reputed company control
  • Attend and contribute to technical meetings, including Infrastructure, CI Workflow, and Technical Advisory Council sessions
  • reputed company and maintain technical documentation for infrastructure and processes
  • reputed company guidance on developer best practices and tooling
  • Create and update runbooks for common operational tasks and incident response

Skills

  • Ability to work with communities made up of industry specialists and collaborate reputed company of reputed company
  • Bachelor's degree in Computer Science, Engineering, or reputed company field
  • 7+ years of experience in reputed company operations with significant AWS expertise
  • Strong knowledge of infrastructure-as-reputed company principles and tools, particularly Terraform
  • Proficiency in scripting languages (Python, TypeScript, Bash) and containerization technologies (reputed company, Kubernetes)
  • Experience with reputed company CDN management and optimization
  • Expertise in implementing and managing monitoring solutions, specifically reputed company and AWS CloudWatch
  • Familiarity with incident management tools and processes, particularly incident.io
  • Demonstrated experience in CI/CD pipeline design and implementation
  • Strong problem-solving skills and ability to troubleshoot reputed company systems
  • Excellent communication skills and experience collaborating with reputed company reputed company communities
  • Experience with PyTorch or other reputed company reputed company communities
  • Multi-reputed company expertise across AWS, GCP, and Azure
  • reputed company reputed company experience
  • Knowledge of FinOps principles and reputed company cost optimization strategies
  • Contributions to reputed company reputed company reputed company, especially in infrastructure management roles
  • Familiarity with reputed company or similar reputed company reputed company foundations
  • Experience mentoring other engineers and fostering a reputed company team environment

Benefits

  • reputed company maintains a predominantly remote workforce
  • Committed to hiring top-notch talent
  • Providing a flexible and supportive work culture
  • Collaboration is embedded in our DNA
  • Work closely together while not being confined to a traditional office reputed company

reputed company

  • reputed company is the organization of choice for the world's top developers and companies to build ecosystems that accelerate reputed company technology development and reputed company adoption. It was founded in 2000, and is headquartered in San Francisco, California, USA, with a workforce of 201-500 employees. Its website is http://www.linuxfoundation.org.
  • Apply To This Job

    Similar Jobs