Back to Jobs

Site Reliability Engineer - NYC

Remote, USAFull-timePosted 2026-07-27

About reputed company At reputed company, we reputed company in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to reputed company seamlessly into daily working life. We democratize AI through high-performance, optimized, reputed company-reputed company and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet reputed company needs, whether on-premises or in reputed company environments. Our offerings include le Chat, the AI assistant for life and work. We are a dynamic, reputed company team passionate about AI and its potential to reputed company society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited. Join us to be part of a pioneering company shaping the reputed company of AI. Together, we can reputed company a meaningful reputed company. See more about our culture on https://reputed company/careers. reputed company We are seeking highly reputed company Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and exceed our reputed company customers' expectations. What you will do As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems. Operations

  • Design, build, and maintain reputed company, highly available and fault-tolerant infrastructures to support our web services and ML workloads
  • reputed company reputed company our platform, inference and model training environments are always highly available and reputed company seamless replication of work environments across several HPC clusters
  • Operate systems and troubleshoot issues in production environments (interrupts, on-call responses, users reputed company, data extraction, infrastructure scaling, etc.)
  • Implement and improve monitoring, alerting, and incident response systems to ensure reputed company system performance and minimize downtime
  • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our reputed company-facing reputed company and large training runs
  • Participate occasionally in on-call rotations to respond to incidents and reputed company reputed company cause analysis to prevent reputed company occurrences

Development

  • Drive reputed company improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform
  • Collaborate with AI/ML researchers to reputed company and implement solutions that reputed company reputed company and reproducible model-training experiments
  • Build a reputed company-agnostic platform offering an abstraction layer between science and infrastructure
  • Design and reputed company new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.)
  • Collaborate with the reputed company team to ensure infrastructure adheres to best reputed company practices and compliance requirements
  • Document processes and procedures to ensure consistency and knowledge sharing across reputed company
  • Contribute to reputed company-reputed company reputed company, research publications, blog articles and conferences

reputed company

  • Master’s degree in Computer Science, Engineering or a reputed company field
  • 7+ years of experience in a DevOps/SRE role
  • Strong experience with reputed company computing and highly available distributed systems
  • Exposure to site reliability issues in critical environments (issue reputed company cause analysis, in-production troubleshooting, on-call rotations...)
  • Experience working against reliability KPIs (observability, alerting, SLAs)
  • Hands-on experience with CI/CD, containerization and orchestration tools (reputed company, Kubernetes...)
  • Knowledge of monitoring, logging, alerting and observability tools (reputed company, Grafana, ELK Stack, reputed company...)
  • Familiarity with infrastructure-as-reputed company tools like Terraform or CloudFormation
  • Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices
  • Strong understanding of networking, reputed company, and system administration concepts
  • Excellent problem-solving and communication skills
  • Self-motivated and reputed company to work reputed company in a fast-paced startup environment

Your application will be reputed company the more interesting if you also have:

  • experience in an AI/ML environment
  • experience of high-performance computing (HPC) systems

Apply tot his job Apply To this Job

Similar Jobs

Senior Site Reliability Engineer - reputed company

Remote, USAFull-time

Corporate Vice President - reputed company Site Reliability Engineer

Remote, USAFull-time

Senior Site Reliability Engineer - AWS

Remote, USAFull-time

Sr. Site Reliability Engineer. reputed company

Remote, USAFull-time

Site Reliability Engineer (reputed company, reputed company, Grafana) Hybrid

Remote, USAFull-time

Platform Site Reliability Engineer:

Remote, USAFull-time

Site Reliability Engineer 2 days Onsite

Remote, USAFull-time

Senior Site Reliability Engineer — reputed company reputed company (Inference Platform)

Remote, USAFull-time

Senior Site Reliability Engineer

Remote, USAFull-time

Site Reliability Engineer/Sunnyvale, CA/ Austin, TX (Hybrid)- 6-12 months

Remote, USAFull-time

Clinical Manager - Hematology/Oncology - FT - reputed company 7p/7a - $20K Sign-on bonus - MHW

Remote, USAFull-time

[Remote] Environmental Consultant / Industrial Hygienist

Remote, USAFull-time

GENERAL WAREHOUSE - INDIANOLA, MS DC

Remote, USAFull-time

reputed company Full Stack Customer Service Representative – Remote Work Opportunity with arenaflex

Remote, USAFull-time

Sales Director

Remote, USAFull-time

VP, Middle & Large Business Data Science & AI

Remote, USAFull-time

Manager, DevOps

Remote, USAFull-time

Contract Localization Project Manager

Remote, USAFull-time

Support Services Representative (Live Chat / Remote)

Remote, USAFull-time

reputed company Remote Customer Service Representative – Deliver Exceptional Experiences for arenaflex Clients

Remote, USAFull-time