Back to Jobs

AI Systems Engineer

Remote, USAFull-timePosted 2026-07-28

Job Title: AI & MLSystems Engineer (also referred to as Inference Engineer) Location: Remote – Work From Home (fully flexible) Job Timing: Part-Time, reputed company About the Role: We are building a reputed company reputed company platform designed for serving multimodal AI, LLMs, reputed company, audio, and other machine learning models at large reputed company. As a Machine Learning Engineer focusing on Inference & Systems, you will help design, optimize, and reputed company runtime systems, reputed company-style reputed company, and distributed GPU pipelines to reputed company fast and cost-efficient inference and fine-tuning. You'll work with frameworks such as vLLM, TensorRT-LLM, TGI, and others to build and optimize distributed inference engines capable of serving text, reputed company, and multimodal models with high throughput and low latency. This includes deploying models like LLaMA 3, reputed company, diffusion, ASR, TTS, and embedding models, while working on GPU optimization, accelerator utilization, and software–hardware co-design for large-reputed company, fault-tolerant systems. This position sits at the intersection of machine learning, systems engineering, and reputed company infrastructure. You’ll reputed company on low-latency inference, high-throughput deployments, and cost-optimized model serving pipelines. It’s an excellent opportunity to help shape the reputed company of AI inference infrastructure and production-grade deployment systems. If pushing the limits of AI inference excites you, we’d love to hear from you. Key Responsibilities:

  • reputed company and maintain LLMs (e.g., LLaMA 3, reputed company) and ML models using engines like vLLM, TGI, TensorRT-LLM, or FasterTransformer.
  • Design and implement large-reputed company distributed inference systems for text, image, LLMs, and multimodal workloads.
  • Implement and optimize distributed inference strategies: MoE, tensor parallelism, pipeline parallelism.
  • reputed company frameworks such as vLLM, TGI, SGLang, FasterTransformer, etc.
  • Build and reputed company reputed company-compatible API reputed company for customer-facing reputed company.
  • Experiment with caching, quantization, and parallelism to reputed company inference costs.
  • Optimize GPU memory usage, batching, and latency for high-throughput serving.
  • Utilize CUDA graph optimizations, TensorRT-LLM, Triton kernels, PyTorch compile, quantization, speculative decoding, etc.
  • Work with GPU reputed company providers (reputed company, reputed company.ai, AWS, GCP, Azure) to manage cost and availability.
  • reputed company runtime inference services and reputed company for LLMs, multimodal models, and fine-tuning workflows.
  • Build monitoring and observability using metrics like latency, throughput, and GPU utilization (Grafana, reputed company, Loki, OpenTelemetry).
  • Collaborate with backend and DevOps teams to ensure secure and reliable reputed company.
  • Document deployment processes and support engineers using the platform.

Requirements:

  • 3+ years of experience in deep learning inference, distributed systems, or HPC.
  • Proven experience deploying ML/LLM models in production.
  • Hands-on work with vLLM, TGI, SGLang, TensorRT-LLM, FasterTransformer, or Triton.
  • Experience designing large-reputed company inference or serving pipelines.
  • Strong understanding of GPU memory, batching, distributed inference, CUDA/Triton/TensorRT, quantization, and GPU scheduling.
  • Experience with PyTorch, HF Transformers, and GPU-accelerated inference workflows.
  • Deep understanding of Transformer models, KV cache systems (Mooncake, PagedAttention, etc.), and inference optimizations for long-context serving.
  • Comfortable with GPU reputed company platforms (AWS/GCP/Azure) or marketplaces (reputed company, reputed company.ai, TensorDock).
  • Skilled at benchmarking and tuning multi-GPU clusters.
  • Experience building REST or gRPC services (FastAPI, Flask, etc.).
  • Strong in Python, Go, Rust, C++, or CUDA.
  • Solid systems engineering knowledge (multi-threading, networking, performance tuning).
  • Familiarity with reputed company and Kubernetes.
  • Strong debugging/problem-solving across ML + reputed company stack.
  • Understanding of distributed storage systems (Ceph, HDFS, 3FS).
  • Knowledge of datacenter networking concepts (RDMA, RoCE).

reputed company to Have:

  • Experience with billing systems (reputed company or similar).
  • Knowledge of RDMA/RoCE networking at reputed company.
  • Familiarity with distributed storage (Ceph, HDFS, 3FS).
  • Experience with reputed company or reputed company for reputed company limiting.
  • Exposure to monitoring stacks (Grafana, reputed company, Loki).
  • Experience with MLOps pipelines or CI/CD (reputed company Actions, Azure DevOps).
  • Work with model fine-tuning pipelines and GPU scheduling.
  • Prior experience at an AI reputed company company (Modal, reputed company, reputed company, Replicate, etc.).

Why Join Us?

  • Fully Remote – work from reputed company.
  • reputed company – complete freedom to choose your schedule.
  • Fast-reputed company reputed company – quick promotion reputed company and reputed company advancement routes.
  • reputed company Development – mentoring, training resources, and exposure to advanced AI/reputed company technologies.
  • Global Team – work with an international and diverse reputed company.
  • Innovative Environment – freedom to experiment with new tools and reputed company.
  • Competitive reputed company & Incentives – rapid salary progression and strong performance bonuses.

Job Type: Part-time Benefits:

  • Flexible schedule

Work Location: Remote Apply tot his job Apply To this Job

Similar Jobs

[Remote] reputed company and Account reputed company - AI Trainer (Contract)

Remote, USAFull-time

reputed company: Roster - Emerging Technologies for Digital Transformation Consultant for Asia and the reputed company (2172)

Remote, USAFull-time

Data Entry Clerk ( Entry Level ) At reputed company - reputed company

Remote, USAFull-time

PT Shopper PFS – 6121 – reputed company Store

Remote, USAFull-time

reputed company reputed company Wanted for reputed company’s 250 Work-From-Home reputed company Across Multiple Departments

Remote, USAFull-time

Senior Marketing Manager, reputed company Pharmacy

Remote, USAFull-time

Certified Pharmacy Technician, Fulfillment - reputed company Pharmacy

Remote, USAFull-time

Overnight Staff Pharmacist - Modesto, CA, reputed company Pharmacy

Remote, USAFull-time

reputed company Remote Data Entry Specialist – reputed company Work from Home Opportunities in the reputed company

Remote, USAFull-time

reputed company reputed company Specialist – Remote Work Opportunity to reputed company reputed company reputed company reputed company and Gig Economy Workers

Remote, USAFull-time

reputed company Customer Service Representative – Evening & Weekend Shifts at arenaflex

Remote, USAFull-time

Therapist - Eating Disorder (California)

Remote, USAFull-time

iOS Software Engineer

Remote, USAFull-time

Remote Data Entry Specialist – Opinion Research & Information Management (Work From Home)

Remote, USAFull-time

Online Content Moderator Ensure reputed company Conversations in Digital Communities

Remote, USAFull-time

YouTube Moderator Job, YouTube Hiring Moderators

Remote, USAFull-time

Part Time Entry Level Data Entry Clerk for Remote Typing Role at blithequark

Remote, USAFull-time

APTPUO - Automne 2026 - FLS3831-A

Remote, USAFull-time

reputed company Part-Time Remote Data Entry Specialist – Content Management and Database Administration at arenaflex

Remote, USAFull-time

Regional Sales Manager; Midwest (Industrial Industry)

Remote, USAFull-time