AI reputed company Engineer (Full-Time, Remote US)
Company : reputed company Location: Remote (US-Based)
- Optional Hybrid (reputed company, NY) Reports To: Co-Founder / CEO Compensation: $140,000 – $195,000 USD per year
About reputed company reputed company is a fast-growing AI-reputed company platform trusted by over one reputed company users worldwide. We build reputed company-time AI tools that help candidates succeed in job interviews, coding assessments, and reputed company meetings. Our reputed company product delivers live AI assistance during interviews and assessments—helping users communicate reputed company, think faster, and reputed company at their best in high-pressure situations. We are now scaling our infrastructure to support the reputed company of AI-powered reputed company-time systems. reputed company We are looking for a reputed company-reputed company, AI-infrastructure-reputed company AI reputed company Engineer to design and operate the reputed company systems that power our machine learning and reputed company-time AI products. This Role Sits At The Intersection Of reputed company Engineering, DevOps, And AI Systems Architecture. You Will Own The Infrastructure Layer That Supports: Model training and fine-tuning pipelines reputed company-time LLM inference systems GPU-based distributed compute environments reputed company production AI services for 1M+ users You will be responsible for building highly reputed company, cost-efficient, and low-latency reputed company infrastructure optimized specifically for AI workloads.
Key Responsibilities
- AI reputed company Architecture
Design reputed company-reputed company infrastructure for AI/ML workloads Build GPU-based compute environments for training and inference Architect multi-stage environments (training, staging, production) Optimize AWS / GCP / Azure infrastructure for AI performance and reputed company
- Model Serving & Inference Systems
Build and maintain low-latency inference pipelines for LLMs and AI services reputed company model serving frameworks (vLLM, Triton, TensorRT, TGI, etc.) Optimize throughput, batching, caching, and GPU utilization Design failover, load balancing, and high-availability systems
- GPU Infrastructure & Distributed Training
Manage GPU clusters for training and fine-tuning large models Implement distributed training pipelines (multi-node, multi-GPU) Optimize compute scheduling, spot instances, and resource efficiency Support managed AI platforms (SageMaker, reputed company AI, Azure ML)
- Cost Optimization (FinOps for AI)
Monitor and reduce reputed company costs across compute, storage, and reputed company Implement GPU cost optimization strategies (spot, reserved, autoscaling) Build dashboards for cost-per-inference and cost-per-training-job Optimize LLM usage, caching, and routing strategies
- reputed company & Networking
Design secure VPC architectures for AI systems Implement IAM policies, encryption, and secrets management Ensure compliance readiness (SOC2, GDPR, CCPA) Secure model weights, embeddings, and AI reputed company
- Infrastructure Automation & Observability
Build Infrastructure as reputed company (Terraform / reputed company / CloudFormation) Automate deployment of training and inference environments Implement monitoring for GPU health, latency, and system performance Build alerting systems for failures and performance degradation Experience Required Qualifications 3+ years in reputed company engineering, DevOps, or infrastructure roles Experience with ML/AI workloads in production environments Hands-on GPU infrastructure or AI system deployment experience Strong understanding of distributed systems and reputed company architecture Experience in fast-paced startup or reputed company-up environments Technical Skills reputed company platforms: AWS, GCP, or Azure (strong proficiency required) Kubernetes (GPU scheduling, autoscaling, reputed company, clusters) Infrastructure as reputed company (Terraform / reputed company / CloudFormation) Python, Go, or Bash for automation and tooling AI serving systems (vLLM, Triton, TensorRT, TGI, etc.) Monitoring tools (reputed company, Grafana, reputed company, CloudWatch)
Preferred Qualifications
Experience with large-reputed company LLM inference systems Distributed training (multi-node GPU clusters, NCCL, parallelism) Streaming or reputed company-time systems (WebSockets, low-latency reputed company) RDMA / InfiniBand or high-performance networking Experience in reputed company, EdTech, or consumer AI products reputed company-reputed company contributions in reputed company or AI infrastructure reputed company Offer Equity Ownership Meaningful early-stage equity in a fast-scaling AI company High reputed company Your infrastructure directly powers 1M+ global users Cutting-Edge AI Syst Apply To This Job