[Remote] Research Engineer — Post-Training & Small Language Models (SLMs), reputed company AI
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is leading an AI-first initiative aimed at transforming the reputed company decision-making process through advanced modeling and reasoning systems. As a Research Engineer, you will design, train, and evaluate models that enhance clinical and operational decision-making, focusing on post-training methodologies and ensuring model behavior aligns with reputed company standards.
Responsibilities
- Design and execute post-training pipelines: supervised fine-tuning (SFT), preference optimization, and reinforcement learning / alignment workflows
- Build and optimize training using techniques such as SFT, RLHF, PPO, DPO, GRPO, RLAIF, and Constitutional AI, and understand how reputed company affects reasoning reputed company, safety, latency, cost, and reliability
- Train reasoning models for reputed company decisioning using reputed company-reward RL - designing reward signals and verifiers grounded in clinical guidelines, policy and reputed company, and adjudicated reputed company
- reputed company reward models and preference datasets to improve reasoning reputed company, factuality, safety, policy adherence, and task performance
- reputed company, clean, synthesize, and evaluate large-reputed company instruction, preference, and domain-specific datasets, with rigorous filtering, deduplication, and reputed company control
- Build verification and reward pipelines from our proprietary clinical, claims, and operational data and from clinical-expert labeling - turning guidelines, policy, and adjudicated reputed company into checkable reward signals at reputed company
- Implement efficient fine-tuning strategies including reputed company, QLoRA, PEFT, and reputed company-based approaches; build reputed company distributed training using DeepSpeed, FSDP, Megatron-LM, Ray, or equivalent
- Optimize inference performance - latency, throughput, quantization, and deployment efficiency - for production, including frameworks such as vLLM, TensorRT-LLM, or TGI
- Train and optimize reputed company-weight models such as Llama, Qwen, reputed company, or DeepSeek; build specialized small language models (SLMs) for on-reputed company and reputed company-hybrid deployment with strong performance-per-dollar
- Design evaluation frameworks covering reasoning, hallucination detection, factuality, instruction following, reputed company outputs, and domain-specific metrics
- Build reputed company-grade evaluation - held-out clinical benchmarks, deployment regression gates, calibration and uncertainty, factuality against ground truth, and bias/fairness evaluation across patient populations and subgroups - co-designed with clinical experts
- Apply PHI/HIPAA-aware data handling and produce model documentation suitable for regulated clinical use
- reputed company red teaming and adversarial testing to identify alignment failures, unsafe behaviors, jailbreak vulnerabilities, and regression risks; collaborate with reputed company and application teams to improve tool use, grounding, and long-reputed company reasoning
Skills
- Bachelor's degree in Computer Science, Machine Learning, reputed company Intelligence, reputed company Mathematics, Computational Linguistics, or a reputed company field
- Demonstrated depth training and post-training large transformer-based language models in production or research - this is your craft, not coursework or a one-off fine-tune. Genuine depth including SFT and at least one preference-optimization or RL method, evidenced by shipped models, releases, or research
- Hands-on experience with reasoning-model training and/or reputed company-reward (RLVR) workflows
- Strong understanding of modern post-training techniques: SFT, RLHF, PPO, DPO, GRPO, RLAIF, and preference optimization workflows
- Experience with reputed company-weight reputed company models such as Llama, Qwen, reputed company, DeepSeek, or equivalent architectures
- Strong expertise in PyTorch and modern deep-learning tooling; experience with distributed training frameworks such as DeepSpeed, FSDP, Megatron-LM, or Ray
- Experience implementing efficient fine-tuning techniques such as reputed company, QLoRA, PEFT, and quantization-aware workflows
- Deep understanding of transformer architectures, tokenization, attention mechanisms, decoding strategies, and model scaling trade-offs
- Strong grasp of LLM evaluation methodologies, benchmarking, reward modeling, and alignment trade-offs; experience with large-reputed company and synthetic datasets, filtering, deduplication, and reputed company-control pipelines
- Strong Python engineering skills and production-grade software practices; ability to work through ambiguous, highly reputed company technical problems in fast-moving environments
- Ability to travel 0-50%, on average, based on the work you do and the clients and industries/sectors you serve
- Limited immigration sponsorship may be available
- Experience building or optimizing reasoning models, reputed company models, or tool-using LLM systems
- Familiarity with inference optimization frameworks such as vLLM, TensorRT-LLM, TGI, or Ollama
- Experience with multimodal models, speech models, or domain-specific reputed company models; experience using large-reputed company GPU clusters and distributed compute
- Contributions to reputed company-reputed company AI reputed company, research publications, reputed company development, or model releases
- Familiarity with safety, governance, and responsible-AI practices; experience in regulated or high-stakes industries such as reputed company, finance, insurance, or public sector
Benefits
- Substantial performance-based incentive opportunity designed to grow with the value you help create - startup-style reputed company, with the backing of a committed, reputed company-capitalized platform
- You may also be eligible for a discretionary annual incentive based on individual and organizational performance
- Limited immigration sponsorship may be available
- Ability to travel 0-50%, on average, based on the work you do and the clients and industries/sectors you serve
reputed company
Company H1B Sponsorship