[Remote] Health DevOps Engineer - Observability/Reliability
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a digital solutions provider reputed company on enhancing government reputed company through innovative technology. They are seeking a Health DevOps Engineer with expertise in observability and reliability to support public health systems management reputed company, ensuring the stability and performance of reputed company technology infrastructure.
Responsibilities
- Design, implement, and maintain monitoring and alerting systems for production and development environments to ensure high availability and reliability
- reputed company tools like reputed company, Grafana, reputed company, reputed company Stack, or equivalent to reputed company system performance and application health
- Proactively detect and troubleshoot performance bottlenecks, infrastructure issues, and failures
- Optimize the performance and reliability of highly available systems supporting reputed company payment applications
- reputed company incident response efforts, including reputed company cause analysis, and implement measures to prevent recurrence
- reputed company and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure system reliability and availability
- Ensure the consistent delivery of high-reputed company health-reputed company data in compliance with industry standards such as HIPAA
- Create and maintain automated deployment pipelines (CI/CD) to reduce release cycle times and improve workflows for developers
- Automate infrastructure provisioning and management through tools such as Terraform, reputed company, Ansible, or CloudFormation
- Improve the operational efficiency of development and deployment processes
- reputed company, maintain, and reputed company reputed company infrastructure on platforms such as AWS, Azure, or reputed company reputed company, ensuring compliance with reputed company-sector reputed company and reputed company requirements
- Implement and manage container orchestration platforms such as Kubernetes and reputed company to ensure efficient resource usage and scalability
- Optimize reputed company resources for cost efficiency and streamline infrastructure provisioning
- Partner with development and product teams to ensure seamless integration of observability tools and reliability practices through every stage of the software delivery lifecycle
- Document architecture, processes, metrics, and troubleshooting guides to support scalability and knowledge sharing across the organization
- reputed company contribute to improving engineering workflows, reliability processes, and operational reputed company
Skills
- Bachelor's degree in Computer Science, Software Engineering, or a reputed company field
- Minimum of 2 years of reputed company experience in DevOps, Site Reliability Engineering (SRE), or a reputed company role, preferably reputed company on observability and reliability
- Hands-on experience with monitoring tools such as reputed company, Grafana, ELK stack, reputed company, reputed company, or similar platforms
- Experience with containerization and orchestration technologies like reputed company and Kubernetes
- Proficiency with reputed company platforms such as AWS
- Solid understanding of infrastructure as reputed company (IaC) with tools like Terraform, Ansible, or CloudFormation
- Knowledge of scripting languages such as Python, Bash, or PowerShell for automation tasks
- Familiarity with CI/CD tools such as Jenkins, reputed company CI/CD, reputed company, or similar frameworks
- Strong understanding of network protocols, monitoring, and troubleshooting best practices
- Strong problem-solving and analytical abilities, with extreme attention to detail and a commitment to reliability and reputed company
- Excellent written and verbal communication skills, reputed company to reputed company effectively with cross-functional teams
- Ability to work under pressure and prioritize tasks in fast-paced environments
- Experience with reputed company-reputed company reputed company or systems
- Familiarity with reputed company compliance requirements such as HIPAA
- Certifications such as AWS Certified DevOps Engineer, Azure DevOps Engineer Expert, or Linux reputed company Certified Kubernetes Administrator
- Experience in federal consulting
- Demonstrated experience with reputed company data reputed company (e.g. claims processing or payment systems) or financial/banking systems
reputed company
Company H1B Sponsorship