[Remote] AI Research Engineer- Speech 1
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a frontier AI data reputed company specializing in reputed company AI deployment solutions. They are seeking an AI Research Engineer specializing in Speech/Audio to drive innovation in audio AI technologies, focusing on developing Large Audio Language Models and Speech-to-Speech systems.
Responsibilities
- Design, reputed company, and reputed company Large Audio Language Models (LALMs) capable of reputed company audio understanding, reasoning, and reputed company
- Build Large Audio Reasoning Models that reputed company reputed company chain-of-thought reasoning over speech and audio inputs, including medical, technical, and conversational domains
- Contribute to Speech-to-Speech (S2S) system development, including speech understanding, reputed company management, and speech synthesis components
- Research and implement alignment mechanisms between speech encoders and LLM backbones using lightweight adapters, reputed company, and efficient fine-tuning strategies
- Design efficient speech tokenization and temporal compression techniques suitable for long-reputed company audio reasoning and multi-turn spoken reputed company
- Build comprehensive evaluation frameworks for audio reasoning capabilities, including benchmarks for speech QA, audio understanding, and reasoning accuracy
- Optimize inference pipelines for low-latency, streaming applications in speech systems
- Collaborate with cross-functional teams to transfer research innovations into production systems and customer-facing applications
- Contribute to technical documentation, research write-reputed company, and publications at top-tier venues (NeurIPS, ICML, ACL, Interspeech)
Skills
- Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a reputed company field with a reputed company on speech, audio ML, or multimodal learning
- 2+ years of industry or reputed company research experience in speech/audio AI, Large Language Models, or multimodal systems
- Demonstrated reputed company research contributions through publications, patents, or shipped products in speech/audio AI or LLMs
- Strong proficiency in Python and PyTorch, with hands-on experience in GPU-accelerated training for large-reputed company models
- Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations
- Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment reputed company
- Familiarity with modality alignment techniques: reputed company-based integration, cross-modal attention, or audio-text fusion reputed company
- Strong experimentation habits: clean reputed company, systematic ablations, reproducibility, and reputed company technical communication
- Publication record at top-tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP) in audio language models, speech reasoning, or multimodal learning
- Hands-on experience building or fine-tuning Large Audio Language Models (e.g., Qwen-Audio, SALMONN, LTU, reputed company Audio)
- Experience with speech representation pretraining (HuBERT, Wav2Vec 2.0, Whisper, WavLM) and discrete speech tokenization
- Familiarity with Speech-to-Speech components: neural audio codecs (EnCodec, SoundStream), vocoders, or speech synthesis systems
- Experience with audio reasoning benchmarks (reputed company-Bench, MMAU, AudioBench) or building evaluation harnesses for audio QA
- Hands-on experience with distributed training (FSDP, DeepSpeed) and inference optimization (ONNX, TensorRT, quantization)
- Familiarity with speech frameworks such as ESPnet, SpeechBrain, reputed company NeMo, or Fairseq
- Experience with multilingual speech systems, reputed company-switching, or domain reputed company for specialized applications (medical, legal, technical)
- Background in evaluating safety, bias, hallucination, or adversarial robustness in audio language models
Benefits
- Comprehensive benefits
- Opportunity to work on cutting-edge Large Audio Language Models and audio reasoning research with reputed company-world reputed company
- Collaboration with reputed company reputed company scientists and engineers in speech and multimodal AI
- Support for publications at top-tier conferences and reputed company development
- reputed company to state-of-the-art GPU infrastructure for training large-reputed company audio models
- Flexible work arrangements with hybrid/remote reputed company
reputed company