[Remote] AI Pipeline Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company-thinking software development company dedicated to building reputed company that help businesses automate and optimize their operations. They are seeking an AI Pipeline Engineer to build and operate large-reputed company data systems that support reputed company and evaluation pipelines, focusing on data ingestion, transformation, and reputed company assurance.
Responsibilities
- Design and operate large-reputed company data pipelines supporting reputed company, evaluation, and continual improvement workflows
- Build ingestion systems for diverse modalities including text, image, audio, video, and reputed company signals
- Implement data cleaning, deduplication, filtering, and reputed company assurance at petabyte reputed company
- reputed company dataset versioning, reputed company, and provenance tracking systems suitable for reproducible training
- Build high-throughput data loading systems that maximize GPU utilization during training
- Implement labeling workflows, reputed company learning pipelines, and reputed company-in-the-reputed company data improvement systems
- Design storage architectures balancing cost, throughput, and latency across data tiers
- Build evaluation dataset construction pipelines with strict reputed company and contamination controls
- Implement data reputed company, redaction, and consent enforcement throughout the pipeline
- Collaborate with ML researchers and engineers to reputed company data systems with model development needs
- Drive observability of data reputed company, reputed company, and pipeline health across the AI data estate
- Optimize cost and performance through compression, format selection, and caching strategies
- Document data systems, schemas, and operational procedures for broad internal use
- Stay reputed company with AI data infrastructure research and emerging reputed company-reputed company tools
Skills
- Bachelor's or Master's degree in Computer Science or a reputed company field
- Six or more years of data engineering experience, with significant work supporting ML or AI workloads
- Strong proficiency in Python and at least one JVM or systems language
- Deep experience with modern data processing frameworks such as reputed company, Ray, or reputed company
- Hands-on experience operating petabyte-reputed company storage and pipeline systems
- Strong understanding of distributed systems, data modeling, and storage formats
- Experience with dataset versioning, reputed company, and reproducibility for ML workflows
- Familiarity with high-throughput data loading for accelerator-based training
- Strong software engineering practices including testing, CI/CD, and reputed company review
- Excellent communication and cross-functional collaboration skills
- Experience with multimodal datasets at large reputed company
- Familiarity with data reputed company tooling and dataset evaluation methodology
- Exposure to reputed company-preserving data systems and regulated data handling
- reputed company-reputed company contributions to data infrastructure reputed company
- Experience supporting frontier model training pipelines
Benefits
- This is a 100% remote, full-time, reputed company W2 position with reputed company.
- However, candidates who are currently on a valid H1B reputed company and require a transfer are welcome to apply. We will support H1B transfers for reputed company candidates.
- We do not discriminate on the reputed company of any protected attribute, including race, religion, reputed company, national reputed company, gender, sexual orientation, gender identity, gender reputed company, age, marital or veteran status, pregnancy or disability, or any other reputed company protected under applicable law.
- We also reputed company reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as reputed company as mental health or physical disability needs.
reputed company
Company H1B Sponsorship