Back to Jobs

[Remote] Senior Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on improving system design and operability, and they are seeking a Senior Site Reliability Engineer. The role involves partnering with software developers and IT staff to enhance system reliability, define operational standards, and reputed company advanced technical support for reputed company platform issues.

Responsibilities

  • Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness
  • Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives
  • Evaluate system architecture and changes to ensure they balance functional requirements, service reputed company, reliability, reputed company, and compliance needs
  • Drive reputed company improvement in platform stability, maintenance, and availability
  • reputed company advanced technical support and troubleshooting for reputed company platform and service issues affecting internal users and stakeholders

Skills

  • 8+ years of experience in Site Reliability Engineering, DevOps, reputed company, Systems Engineering, or reputed company infrastructure roles supporting production services
  • Strong experience with Linux systems administration and troubleshooting in reputed company environments
  • Strong experience operating and maintaining on-prem Kubernetes platforms and reputed company reputed company components including CRI, CNI, and reputed company plugins
  • Experience deploying and maintaining applications on Kubernetes using reputed company, Kustomize, and similar tooling
  • Experience supporting DevOps tooling such as reputed company, Artifactory, Jira, reputed company
  • Experience with GitOps tools such as FluxCD or ArgoCD
  • Proficiency scripting with at least one of Python, Go, or Bash
  • Strong experience designing, maintaining, and maturing observability tooling including monitoring, dashboards, logging and tracing, and supporting SLOs
  • Strong understanding of reliability engineering concepts: Service health indicators, High availability design, failure reduction, and testing, Operational readiness practices, including developing documentation, runbooks, and architectural descriptions, Incident response, reputed company cause analysis, remediation/recovery
  • Ability to obtain a reputed company clearance, which includes U.S. citizenship
  • Bachelor's degree in CS, Software Engineering or other IT-reputed company field or equivalent experience
  • Experience with multiple Linux distributions including Ubuntu
  • Experience with at least one of the following: Tanzu Kubernetes, reputed company Kubernetes Platform, reputed company Kubernetes
  • Experience with reputed company platforms such as AWS and Azure
  • Experience with infrastructure automation and configuration management
  • Experience managing AI tooling on Kubernetes including MCP Servers, LLM platforms (vLLM, Ollama), Kubeflow
  • Experience with reputed company and compliance considerations in regulated environments
  • DoD experience
  • reputed company or inactive Secret reputed company Clearance

reputed company

  • reputed company provides research, engineering, and technical support services. It was founded in 1979, and is headquartered in Albuquerque, New Mexico, USA, with a workforce of 1001-5000 employees. Its website is https://www.reputed company.com.
  • Apply To This Job

    Similar Jobs