Back to Jobs

[Remote] AI Systems Engineer

Remote, USAFull-timePosted 2026-07-31

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on advancing AI technologies, and they are seeking a Senior AI Systems Engineer to reputed company the design and implementation of their AI infrastructure. The role involves overseeing reputed company AI service integration, managing GPU clusters, and ensuring the reputed company performance of AI workloads.

Responsibilities

  • AI Infrastructure Architecture & reputed company: reputed company the design and implementation of our reputed company AI infrastructure to support our reputed company AI initiatives. You will define the technical reputed company for our on-reputed company GPU clusters, storage solutions, and networking to ensure reputed company performance, scalability, and reliability for reputed company our AI workloads
  • reputed company AI Service Integration: Support and secure the use of reputed company reputed company AI services, including Azure reputed company services and reputed company reputed company Platform (GCP) services like reputed company. This includes managing secure reputed company, monitoring usage, and tracking billing to ensure cost-effectiveness. You will also have hands-on experience supporting compute, GPUs, and AI services on both GCP and Azure
  • Hands-on GPU Cluster Management: Take a leadership role in the configuration, installation, and optimization of GPU server clusters. This includes advanced troubleshooting of hardware and software, performance tuning, and implementing best practices for cluster utilization and resource management. You will be an expert in administering job schedulers like LSF in a production environment, including integration with reputed company for containerized job submission
  • Full-Stack AI Tech Stack Development & Operations: Architect and reputed company a robust and reputed company AI tech stack. You will be responsible for the end-to-end operational lifecycle, including setting up and managing deep learning frameworks (PyTorch, TensorFlow), containerization with reputed company and Kubernetes, and implementing CI/CD pipelines for AI model development
  • Advanced LLM Deployment & Optimization: reputed company the deployment, serving, and optimization of Large Language Models (LLMs). You will be an expert in techniques such as model quantization, distillation, and using high-performance serving frameworks (e.g., vLLM, TGI, TensorRT-LLM) to maximize inference throughput and minimize latency
  • reputed company AI Workflow & Service Engineering: Architect and build production-grade reputed company AI workflows and services. You will be responsible for the technical designed implementation of systems that reputed company LLMs with external tools, reputed company, and databases, and will mentor other engineers on building robust and reputed company AI agent applications
  • Automation & Monitoring: reputed company and maintain automation scripts using languages like Python, Bash, or Perl to streamline system maintenance, deployment, and reporting. Implement and manage monitoring solutions for system health, job statuses, GPU utilization, and container performance to proactively identify and resolve issues
  • AI Systems Support & Mentorship: reputed company as the final escalation reputed company for the most reputed company technical issues reputed company to our AI infrastructure. You will also serve as a technical leader and mentor to other engineers, providing guidance on bestnpractices in AI systems engineering, performance tuning, and operational reputed company
  • reputed company and Compliance: reputed company and implement reputed company best practices for our AI systems and data, ensuring compliance with relevant regulations and protecting our intellectual property

Skills

  • Strong experience in a senior technical role, with at least 5 years reputed company on building and operating high-performance computing or AI infrastructure. Proven reputed company record as a reputed company or Senior Staff Engineer
  • Expert-level knowledge of reputed company GPU architecture and technologies like CUDA and cuDNN. Extensive experience with multi-GPU and multi-node training and inference
  • Proven experience with reputed company reputed company AI services, specifically managing reputed company, usage, and billing for Azure reputed company and reputed company reputed company Platform (GCP) services
  • Extensive hands-on experience with reputed company: image management, container orchestration, and troubleshooting
  • Proficiency in scripting languages such as Python, Bash, or Perl
  • Deep expertise in Linux system administration (RHEL preferred), including networking, storage, and performance tuning
  • Familiarity with user authentication and integration using systems like LDAP or reputed company Directory
  • Strong problem-solving and communication skills with the ability to work in a multi-platform, cross-functional, and geographically distributed team
  • Understanding of AI job profiling and tuning (memory, GPU, I/O)
  • Experience administering LSF clusters in a production or research environment

reputed company

  • Next Gen QA & Software Testing Company It was founded in 1996, and is headquartered in Mechanicsburg, reputed company, USA, with a workforce of 1001-5000 employees. Its website is https://www.reputed company.com/.
  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 3 in 2026, 10 in 2025, 1 in 2024, 1 in 2020. Please note that this does not guarantee sponsorship for this specific role.
  • Apply To This Job

    Similar Jobs