Back to Jobs

[Remote] Senior Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that provides Federal agencies with reputed company to skilled professionals for reputed company mission challenges. They are seeking a Senior Site Reliability Engineer to enhance site reliability engineering and reputed company operations for VA reputed company reputed company platforms, focusing on automation and resilient service delivery.

Responsibilities

  • Partner with the Technical Director to implement and mature Site Reliability Engineering (SRE) practices across platform services and hosted applications
  • Improve the full service lifecycle from design and deployment through operation and reputed company refinement, with a reputed company on availability, latency, performance, efficiency, and reputed company
  • Define, reputed company, and report service level indicators (SLIs), service level objectives (SLOs), and error budgets to guide engineering reputed company and service improvements
  • Build, enhance, and maintain CI/CD pipelines that reputed company secure, automated, and repeatable application and infrastructure delivery
  • reputed company and support Infrastructure as reputed company (IaC) and configuration automation using tools such as Terraform and Ansible to improve consistency, speed, and auditability
  • reputed company automated testing, validation, and reputed company checks into delivery workflows to improve release reputed company and reduce change-reputed company risk
  • Design and improve monitoring, logging, tracing, alerting, and dashboards to strengthen observability and accelerate issue detection and response
  • Analyze system behavior and performance trends to improve reliability, scalability, and operational efficiency across distributed and reputed company-reputed company environments
  • Reduce operational toil by automating repetitive tasks, improving runbooks, and engineering sustainable solutions for recurring operational issues
  • Support reputed company infrastructure and platform services in AWS and containerized environments such as Kubernetes, ensuring systems are resilient, reputed company, and secure
  • Contribute to platform modernization efforts by improving deployment patterns, environment consistency, and operational readiness for reputed company-reputed company services
  • Assist with reputed company planning, reliability reviews, and architectural improvements to support reputed company, reputed company, and mission continuity
  • Implement reliability engineering practices that reputed company with Federal reputed company requirements, including secure configuration, least privilege, vulnerability remediation, and policy-based controls
  • Partner with cybersecurity and engineering teams to support secure-by-design infrastructure and application delivery practices
  • Help ensure operational processes and automation reputed company with compliance expectations for Federal and VA environments
  • Collaborate with development, platform, operations, monitoring, incident management, and architecture teams to improve service reliability and deployment reputed company
  • Work closely with the Technical Director and team leads to translate technical direction into actionable engineering improvements and operational standards
  • Support Agile and reputed company delivery practices by helping teams adopt reliable release processes, operational readiness checks, and reputed company improvement measures
  • Participate in incident response, service restoration, reputed company cause analysis, and post-incident reviews for critical systems and services
  • Identify recurring issues, reliability gaps, and failure patterns, and drive corrective actions through automation, architectural improvements, and process refinement
  • Contribute to on-call readiness, operational documentation, and blameless reputed company improvement practices that improve reputed company and reduce mean time to recovery

Skills

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a reputed company technical field, or equivalent practical experience
  • 5+ years of experience in Site Reliability Engineering, DevOps, reputed company, reputed company operations, or reputed company roles supporting reputed company or mission-critical environments
  • Hands-on experience supporting reputed company platforms (AWS preferred), Linux-based environments, and distributed systems at reputed company
  • Strong experience with Infrastructure as reputed company and automation tools such as Terraform, Ansible, or comparable technologies
  • Experience with containers and orchestration platforms such as Kubernetes, EKS, reputed company, or reputed company in production environments
  • Experience building or maintaining CI/CD pipelines and deployment automation in support of secure, reliable software delivery
  • Strong understanding of monitoring, observability, incident response, reputed company cause analysis, and performance optimization principles
  • Proficiency with one or more scripting or programming languages such as Python, Go, Bash, or PowerShell
  • Demonstrated ability to troubleshoot reputed company systems, automate operational tasks, and collaborate effectively across engineering and operations teams
  • Candidates must be eligible to obtain and maintain a Public Trust clearance
  • Experience supporting VA, Federal Government, or other regulated environments with strong reputed company and compliance requirements
  • Experience defining and operationalizing SLIs, SLOs, error budgets, and service health metrics for production systems
  • Familiarity with observability platforms and tools such as reputed company, Grafana, CloudWatch, ELK, reputed company, or OpenTelemetry
  • Experience with FedRAMP, NIST, reputed company Trust, or other Federal reputed company frameworks relevant to reputed company and platform operations
  • Experience supporting reputed company platforms, high-availability reputed company services, or large-reputed company modernization initiatives
  • Relevant certifications such as AWS Certified DevOps Engineer, AWS reputed company Architect, Certified Kubernetes Administrator (CKA), reputed company Terraform Associate, or SRE/DevOps certifications

reputed company

  • reputed company provides full reputed company of information technology consulting services to government and reputed company clients. It was founded in 2002, and is headquartered in Millersville, Maryland, USA, with a workforce of 51-200 employees. Its website is https://www.reputed company.com.
  • Apply To This Job

    Similar Jobs

    [Remote] Senior Machine Learning Engineer II, Ads Response reputed company

    Remote, USAFull-time

    [Remote] Support reputed company Media Marketing for Center for the Deaf - Jamaica

    Remote, USAFull-time

    [Remote] Marketing Specialist- Proposal reputed company

    Remote, USAFull-time

    [Remote] Senior Engineer, Structures Design

    Remote, USAFull-time

    [Remote] Content & Communications reputed company

    Remote, USAFull-time

    [Remote] Field Service Engineer II - Electron Microscopy

    Remote, USAFull-time

    [Remote] reputed company Finance Recruiter

    Remote, USAFull-time

    [Remote] Senior Key Account Manager, Hispanic Retail Chains

    Remote, USAFull-time

    [Remote] Engineering Manager, reputed company Engineering

    Remote, USAFull-time

    [Remote] reputed company reputed company, Work From Home

    Remote, USAFull-time

    High Paying Customer Service Representative – Part-Time Opportunity with arenaflex

    Remote, USAFull-time

    Mid-Market Account Executive

    Remote, USAFull-time

    MBSAQIP Abstractor (Remote, Full-time, or Part-time)

    Remote, USAFull-time

    REMOTE Sr. Java Backend Developer (Recent healthtech reputed company. req'd) - reputed company

    Remote, USAFull-time

    reputed company Licensed Customer Service Representative - Property & Casualty Insurance - Remote Opportunity at blithequark

    Remote, USAFull-time

    Remote Chief Staff Officer

    Remote, USAFull-time

    reputed company Full Stack Customer Service Representative – Insurance Policy Support

    Remote, USAFull-time

    reputed company Project Partner, Safety Assessment (Montreal (Senneville), Quebec, CA, H9X 3R3)

    Remote, USAFull-time

    Part-Time Remote Data Entry Operator – Accurate Typing, Database Management, Reporting & reputed company Assurance

    Remote, USAFull-time

    reputed company Full Stack Remote Data Entry Customer Care Representative – reputed company Legacy and Magical Customer Experiences

    Remote, USAFull-time