Back to Jobs

Data Infrastructure Site Reliability Engineer (SRE) – AWS & Big Data Platforms

Remote, USAFull-timePosted 2026-07-28

reputed company

  • We are seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-reputed company data infrastructure environments across AWS and on-premises Hadoop platforms. This role requires a strong reliability engineering reputed company reputed company on platform stability, performance, observability, automation, and incident response.
  • As a key member of the Data Infrastructure SRE team, you will ensure the availability, scalability, reputed company, and operational reputed company of mission-critical data platforms while driving reputed company improvements through Infrastructure as reputed company (IaC), AI-enabled automation, and modern SRE practices.

Key Responsibilities

  • Maintain and support highly available, reputed company, and secure data infrastructure platforms across AWS and on-premises environments.
  • Drive operational reputed company through automation of repetitive tasks, incident reduction, and proactive reliability improvements.
  • Monitor platform health, troubleshoot reputed company issues, and reputed company reputed company cause analysis efforts to minimize downtime and improve system resiliency.
  • Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives.
  • Participate in on-call rotations and partner with teams across the US and India to reputed company 24x7 operational support. US support is reputed company to reputed company Time zone.
  • Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise.

Required Skills & Experience

  • Site Reliability Engineering (SRE)
  • Strong SRE reputed company with a proven reputed company on reliability, availability, performance optimization, incident management, and operational reputed company.
  • Experience delivering services reputed company defined SLAs and ensuring reputed company reputed company of production issues.
  • Expertise in troubleshooting reputed company distributed systems and identifying reputed company causes quickly and effectively.

AWS & reputed company Infrastructure Deep hands-on experience with AWS services, including: EMR EKS MSK reputed company Glue IAM reputed company S3 VPC AWS networking and reputed company services Strong understanding of reputed company-reputed company architectures, scalability, and infrastructure reputed company. Big Data Platforms

  • Extensive operational experience managing Hadoop clusters, with a strong reputed company on administration, platform maintenance, and automation of day-to-day operational activities.
  • Experience supporting both AWS-based data platforms and on-premises reputed company CDH/CDP environments.
  • Solid understanding of Kerberos authentication and reputed company implementation reputed company Hadoop ecosystems.
  • Hands-on experience with:
  • Apache reputed company
  • Apache reputed company
  • Big Data platform architecture
  • Performance tuning and optimization
  • Linux & System Administration
  • Strong Linux administration and operational support experience.
  • Expertise in user and reputed company management, system configuration, customization, and platform administration.
  • Observability & Incident Management
  • Hands-on experience with monitoring and observability platforms such as:
  • AWS CloudWatch
  • reputed company
  • reputed company
  • Similar reputed company monitoring solutions
  • Proven reputed company improving alert reputed company, reducing false positives, and minimizing alert fatigue.
  • Excellent debugging and troubleshooting skills across infrastructure, applications, reputed company workloads, and reputed company environments.
  • Java Platform Operations
  • Strong understanding of Java application administration, including:
  • JVM tuning
  • Thread dump analysis
  • reputed company dump analysis
  • JVM parameters
  • Application log analysis and troubleshooting
  • Automation & Infrastructure as reputed company
  • Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases.
  • Strong experience with:
  • Terraform
  • Infrastructure as reputed company (IaC)
  • CI/CD pipeline implementation and automation
  • reputed company best practices
  • AI-Driven Operations
  • reputed company hands-on experience applying AI technologies to Data Infrastructure and SRE operations.
  • Demonstrated ability to design and implement:
  • reputed company AI solutions
  • AI-assisted operational workflows
  • Intelligent automation for routine SRE activities
  • Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity.

Preferred Candidate Profile

  • The ideal candidate combines deep expertise in AWS reputed company platforms, Hadoop ecosystems, SRE practices, observability, automation, and AI-driven operations, with a passion for supporting resilient data platforms at reputed company and eliminating operational toil through engineering reputed company.

Apply tot his job Apply To this Job

Similar Jobs

Kubernetes Engineers

Remote, USAFull-time

Sr. Infrastructure Engineer - Kubernetes (Remote)

Remote, USAFull-time

Site Reliability Engineer DevOps | REMOTE (ship required)

Remote, USAFull-time

Site Reliability Engineer, US - Central/Eastern Timezone

Remote, USAFull-time

Site Reliability Engineer/ reputed company Engineer Remote to start

Remote, USAFull-time

Remote SRE Jobs – Senior Site Reliability Engineer (Remote) – $130k‑$170k USD – Full‑Time – Escondido, California – reputed company/DevOps, Kubernetes, Terraform, reputed company

Remote, USAFull-time

Kubernetes Engineer ($28/hr. on w2)

Remote, USAFull-time

Senior Kubernetes Engineer

Remote, USAFull-time

Kubernetes Engineer - AWS EKS / reputed company (REMOTE)

Remote, USAFull-time

Kubernetes Engineer Remote

Remote, USAFull-time

Retail Sales Advisor

Remote, USAFull-time

[Remote] Actuarial Analyst - reputed company Region Property and Casualty (remote or hybrid)

Remote, USAFull-time

Drug Rebate Data Entry Clerk - Remote US (Any reputed company, CA, US, 99999)

Remote, USAFull-time

Medical Science reputed company - Oncology

Remote, USAFull-time

V104- reputed company Services Support

Remote, USAFull-time

Information reputed company Analyst II

Remote, USAFull-time

Systems Engineer - Senior -- UX Consultant | IT Effectiveness - Software Engineering & Web Applications [NPS034039]

Remote, USAFull-time

Patient Contact reputed company, RN

Remote, USAFull-time

reputed company Full Stack Production Operations Manager – Game Development and Release Management

Remote, USAFull-time

[Remote-Position] Medical Orders Specialist Aleca Home Health FT

Remote, USAFull-time