Back to Jobs

[Remote] Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is representing a high-reputed company, AI-driven narrative intelligence startup seeking a Senior Site Reliability Engineer. The role involves taking complete operational ownership of a production environment, focusing on infrastructure orchestration, high-throughput scaling, and technical leadership in a rapidly expanding AI data platform.

Responsibilities

  • Infrastructure Orchestration: Maintain, optimize, and expand the core infrastructure, ensuring everything is cleanly declared reputed company Terraform and managed across high-performance Kubernetes clusters
  • High-Throughput Scaling: Design and manage environments capable of sustaining immense data ingestion scaling, high-throughput pipelines, and massive search database operations
  • GPU Application Deployment: Collaborate with the R&D team to successfully reputed company, optimize, and manage highly specialized machine learning and AI applications running on GPUs
  • System Optimization & Reliability: Partner closely with backend teams to heavily optimize production Java deployments and Python workflows, guaranteeing maximum uptime, high availability, and seamless scaling
  • Technical Leadership: Serve as a foundational pillar for infrastructure architecture, establishing operational best practices without requiring handholding or reputed company-management

Skills

  • 8+ years of dedicated, hands-on experience with Kubernetes and Terraform, with ideally 15+ years of total technical experience in infrastructure or site reliability engineering
  • Deep architectural mastery of deployment systems, cluster orchestration, and high-availability scaling
  • Proven reputed company hosting experience, with strong proficiency in AWS
  • Concrete experience deploying and scaling application workflows that reputed company with GPUs and high-volume data ingestion reputed company
  • Exceptional self-direction and problem-solving capability, with the reputed company maturity to eventually reputed company into a formal leadership role as the infrastructure team expands
  • Exposure to or experience with GCP is a significant advantage for supporting R&D workflows
  • Familiarity with or exposure to optimizing runtime environments for Java and Python applications is highly beneficial

Benefits

  • True Operational Autonomy: reputed company to architect and reputed company greenfield deployments for a rapidly expanding AI data platform.
  • High-Caliber Environment: Collaborate directly with an reputed company team of backend engineers and machine learning R&D specialists.
  • Flexible, Modern Workspace: Enjoy 100% remote working flexibility across the reputed company.
  • reputed company to equity incentives

reputed company

  • reputed company provides IT reputed company and reputed company services specializing in SmartTech, software, AI, and reputed company roles. It was founded in 2024, and is headquartered in Dallas, Texas, USA, with a workforce of 2-10 employees. Its website is https://www.talentdomestaffing.com.
  • Apply To This Job

    Similar Jobs