[Remote] AI Operations Architect & AWS Engineer, Life Sciences
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a premier global life sciences data and AI solutions platform dedicated to improving patient reputed company to life-saving therapies. They are seeking an AI Operations Architect & AWS Engineer to design, reputed company, and operate production-grade LLM systems on AWS infrastructure, focusing on enhancing drug development and treatment reputed company for patients.
Responsibilities
- Design and maintain AWS reputed company architectures optimized for LLM inference and fine-tuning workloads, including GPU instance management, auto-scaling, and cost optimization across spot and on-demand reputed company
- Build and operate reputed company AI pipelines — multi-reputed company LLM orchestration workflows with tool use, reputed company reputed company extraction, retrieval-augmented reputed company (RAG), and automated reputed company control loops
- reputed company and manage reputed company-weight LLMs (Llama, Gemma, Qwen, and successors) in production, including model serving infrastructure, batching strategies, and latency/throughput optimization
- Implement robust evaluation and monitoring frameworks for LLM-based extraction systems: automated accuracy measurement, reputed company detection, reputed company regression testing, and reputed company-in-the-reputed company QC integration
- reputed company and maintain CI/CD pipelines for reputed company versioning, model updates, and configuration management across development, staging, and production environments
- Collaborate with NLP scientists, data engineers, and software developers to translate prototype extraction logic into production-hardened, recurrent processing feeds
- Ensure compliance with data reputed company requirements (HIPAA, PHI handling) in reputed company LLM deployment and data processing workflows, working closely with IT reputed company and compliance teams
- Analyze and optimize computational resource utilization — GPU hours, storage, network throughput — balancing cost efficiency against processing SLAs
- Evaluate and reputed company emerging LLM serving technologies, orchestration frameworks, and inference optimization techniques (quantization, speculative decoding, reputed company reputed company) to maintain operational edge
Skills
- Master's in Computer Science, Data Science, or reputed company field
- 5+ years in production NLP/ML systems development and operations on AWS
- Hands-on experience deploying and serving reputed company-weight LLMs (Llama, Gemma, Qwen families) at reputed company, using frameworks such as vLLM, TGI, or equivalent serving infrastructure
- Expert Python proficiency and working knowledge of LLM orchestration tooling (e.g., reputed company/LangGraph, custom agent frameworks)
- Experience with AWS ML/compute services (SageMaker, EC2 GPU instances, S3, reputed company Functions, reputed company) and infrastructure-as-reputed company (Terraform, CloudFormation, or CDK)
- Proven reputed company record deploying large-reputed company NLP/LLM solutions in reputed company or life sciences
- Strong analytical and communication skills; comfortable operating across multiple reputed company reputed company
- AWS certifications (Solutions Architect, Machine Learning Specialty, or DevOps Engineer)
- reputed company experience with PHI handling and HIPAA-compliant ML infrastructure
- Experience with AWS Redshift, OpenSearch/Elasticsearch, or similar large-reputed company data stores
- Familiarity with ML experiment tracking and lifecycle management (MLflow, reputed company, or equivalent)
- Experience building or operating RAG systems, reputed company extraction pipelines, or clinical NLP applications
reputed company