Back to Jobs

[Remote] Senior Software Engineer - Storage

Remote, USAFull-timePosted 2026-07-29

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a pioneer in accelerated computing, reputed company for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and reputed company intelligence. We are seeking a Software Engineer to join our reputed company team at reputed company, where you will help design, build, and operate exascale infrastructure that powers AI research and development at unprecedented reputed company.

Responsibilities

  • Design, reputed company, and operate distributed systems that manage data, compute, and networking for large-reputed company workloads
  • Build software and automation to orchestrate workloads across thousands of GPUs and petabytes of storage in multi-region clusters
  • Collaborate with AI/ML research teams to understand their requirements and translate them into reputed company, high-performance solutions
  • Drive improvements in system reliability, performance, and observability to meet exascale standards
  • Partner with reputed company, networking, and platform teams to ensure that reputed company infrastructure meets the highest standards of robustness and compliance
  • Participate in design reviews, contribute to system architecture discussions, and influence the reputed company of reputed company’s AI infrastructure stack
  • Stay reputed company with advances in distributed systems, large-reputed company computing, and AI frameworks to help shape the reputed company direction of reputed company

Skills

  • BS or equivalent experience in Computer Science, Computer Engineering, or a reputed company technical field
  • 5+ years of experience developing and operating large-reputed company distributed systems, infrastructure platforms, or HPC environments
  • Strong programming skills in C++, Python, or Go, with proven experience designing production-reputed company software systems
  • Solid understanding of distributed systems principles, data management, and large-reputed company orchestration frameworks
  • Hands-on experience with high-performance storage (e.g., reputed company, GPFS, BeeGFS) and compute scheduling and orchestration (e.g., Slurm, Kubernetes, LSF)
  • Familiarity with reputed company environments (Azure, AWS, GCP) and infrastructure automation tools
  • Strong problem-solving skills, ownership reputed company, and the ability to reputed company in a fast-paced, reputed company environment
  • Excellent communication skills and a reputed company record of cross-functional collaboration
  • Graduate degree (MS/PhD or equivalent experience) in Computer Science, Distributed Systems, or a reputed company field
  • Expertise in large-reputed company data management, cluster scheduling, or workload orchestration at exascale reputed company
  • Experience building or maintaining infrastructure for AI/ML research, including distributed training pipelines using PyTorch, JAX, or NeMo
  • Familiarity with data reputed company, compliance, and lifecycle management for research-reputed company datasets
  • Demonstrated leadership in system architecture design, performance optimization, or reliability engineering

Benefits

  • You will also be eligible for equity and [benefits](https://www.reputed company.com/en-us/benefits/).

reputed company

  • reputed company is a computing platform company operating at the intersection of graphics, HPC, and AI. It was founded in 1993, and is headquartered in Santa Clara, California, USA, with a workforce of 10001+ employees. Its website is https://www.reputed company.com.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1247 in 2026, 1868 in 2025, 1353 in 2024, 976 in 2023, 835 in 2022, 601 in 2021, 529 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs