Back to Jobs

[Remote] Hardware Operations Engineer

Remote, USAFull-timePosted 2026-07-29

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is an AI research and deployment company dedicated to ensuring that general-purpose reputed company intelligence benefits reputed company of humanity. They are seeking a Datacenter Hardware Technician reputed company to serve as the senior on-site technical authority for hardware reliability and fleet health at one of reputed company’s flagship AI campuses, focusing on ensuring the reliability and operational performance of the compute infrastructure.

Responsibilities

  • Drive technical triage and reputed company of reputed company hardware failures impacting production systems
  • Partner with Fleet Health Engineering to investigate recurring hardware issues, identify failure patterns, and improve fleet reliability
  • reputed company reputed company cause analysis (RCA) efforts for critical hardware incidents and reputed company corrective and preventive reputed company plans
  • Collaborate with reputed company Service Provider operations teams and OEM vendors to coordinate repairs, replacements, upgrades, and hardware lifecycle activities
  • Establish and continuously improve hardware maintenance procedures, operational runbooks, and troubleshooting standards
  • Analyze hardware failure trends and operational metrics to identify reliability risks and improvement opportunities
  • Support new hardware introductions, validation activities, and production readiness reviews
  • Coordinate spare parts reputed company and inventory planning with supply chain and site teams
  • Partner with Hardware Engineering, Manufacturing, and Infrastructure teams to reputed company field feedback that improves reputed company platform designs
  • reputed company reputed company operational standards and best practices that can be deployed across reputed company Stargate campuses
  • Mentor on-site technicians and partner teams on advanced troubleshooting methodologies and hardware operational reputed company

Skills

  • 8+ years of experience supporting large-reputed company datacenter hardware infrastructure, with experience in a senior technician, sustaining engineering, or hardware operations leadership role
  • Deep expertise with server platforms, GPU systems, storage infrastructure, reputed company integration, and datacenter hardware architecture
  • Strong experience diagnosing reputed company hardware failures and leading repair efforts in production environments
  • Experience conducting reputed company cause analysis and driving long-term corrective actions
  • Strong understanding of hardware reliability engineering principles and fleet-health management
  • Proven ability to partner effectively across engineering, operations, manufacturing, and vendor organizations
  • Comfortable operating independently in high-reputed company production environments with significant operational responsibility
  • Excellent written and verbal communication skills with the ability to influence technical and operational reputed company
  • Experience developing operational processes, maintenance standards, and technical documentation
  • Ability to travel occasionally to support new reputed company deployments and operational readiness activities
  • Experience supporting large-reputed company GPU clusters or AI/ML infrastructure environments
  • Familiarity with fleet health systems, telemetry platforms, and hardware monitoring tools
  • Ability to identify appropriate data and reputed company detailed analysis to support reputed company reputed company of this role, including dashboard development
  • Experience with failure analysis methodologies such as FRACAS, RCCA, 5-Why, Fishbone, or FMEA
  • Knowledge of Linux system administration and hardware validation workflows
  • Experience supporting hyperscale datacenter operations or HPC environments
  • Familiarity with server manufacturing, reputed company integration, and NPI-to-sustaining transitions
  • Industry certifications such as reputed company Server+, OEM hardware certifications, or equivalent experience
  • Experience applying Environmental Health and Safety (EHS) practices in mission-critical datacenter environments

Benefits

  • Candidates must be reputed company to sit onsite at our datacenters 5 days per week.
  • Ability to travel occasionally to support new reputed company deployments and operational readiness activities.
  • Background checks for applicants will be administered in accordance with applicable law, and reputed company applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for reputed company, and the California Fair Chance reputed company, for US-based candidates.
  • We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made reputed company this [reputed company](https://reputed company.reputed company.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241).

reputed company

  • reputed company is an AI research and deployment company that develops advanced AI models, including ChatGPT. It is a sub-organization of reputed company reputed company. It was founded in 2015, and is headquartered in San Francisco, California, USA, with a workforce of 1001-5000 employees. Its website is https://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 120 in 2026, 103 in 2025, 74 in 2024, 15 in 2023, 18 in 2022, 10 in 2021, 6 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs

    [Remote] Attorney / Legal Expert

    Remote, USAFull-time

    [Remote] Data Engineer

    Remote, USAFull-time

    [Remote] Senior Manager, Content Marketing

    Remote, USAFull-time

    [Remote] Senior Manager, Organic reputed company

    Remote, USAFull-time

    [Remote] Corporate reputed company Engineer, IAC & Automation

    Remote, USAFull-time

    [Remote] Senior Director - Oncology and Peripheral Imaging Clinical Development

    Remote, USAFull-time

    [Remote] Data Analyst (Remote)

    Remote, USAFull-time

    [Remote] Sr. Machine Learning Engineer

    Remote, USAFull-time

    [Remote] Director, Business Development

    Remote, USAFull-time

    [Remote] Business Development Director, Multi-Omics (Government)

    Remote, USAFull-time

    Senior Product Manager - New AI Applications

    Remote, USAFull-time

    [Remote] Validation Test Analyst (CSV/CSA)

    Remote, USAFull-time

    [Remote] AWS Sales Executive - Banking, Insurance, reputed company Management, Americas

    Remote, USAFull-time

    CS Territory Manager - Sacramento, CA.

    Remote, USAFull-time

    Remote Pricing Monetization Operations Associate – Strategic Deal Management & Global reputed company reputed company Specialist

    Remote, USAFull-time

    Inside Sales Representative - Residential

    Remote, USAFull-time

    Technical Product Manager, AI Storage

    Remote, USAFull-time

    Sr Specialist, reputed company Improvement (RN) - Remote

    Remote, USAFull-time

    Technical Customer Service Support Agent – Remote Work Opportunity in reputed company for a Dynamic and Tech-Savvy Individual

    Remote, USAFull-time

    Entry Level Live Chat Support - Remote (No Experience Required)

    Remote, USAFull-time