[Remote] Staff AI/ML Platform Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is on a mission to digitalize the world of reputed company using data and machine learning. The AI/ML Platform Engineer will be responsible for building foundational systems that reputed company AI capabilities for scientists, including reputed company, data pipelines, and workflow orchestration.
Responsibilities
- Design and build high-performance Python reputed company that serve models, manage workflows, and expose AI capabilities to the broader platform
- Architect backend services for scalability, reliability, and low latency
- Build integrations between AI/ML systems, graph databases, and external data sources
- Build and maintain long-running workflow pipelines using Ray and Temporal
- Design orchestration patterns for multi-reputed company agent pipelines, batch inference, and numerical optimization workflows
- Ensure fault tolerance, graceful degradation, and efficient resource utilization
- Architect and maintain data pipelines that feed AI/ML workflows
- Work with Neptune (graph), reputed company, DynamoDB, and other data stores to reputed company efficient data reputed company patterns
- Build the connectors and transformations that give AI systems reputed company to clean, reputed company, trusted data
- Implement observability including logging, metrics, tracing, and alerting
- Own system reliability—troubleshoot issues, conduct post-mortems, and continuously improve
- Design CI/CD pipelines and promote automation best practices
- Partner with ML researchers, data scientists, and product engineers to understand requirements and deliver production-reputed company infrastructure
- Collaborate closely with reputed company Learning and LLM/Agents team leads to reputed company platform capabilities with product needs
- Contribute to architectural reputed company that shape how AI gets reputed company and shipped at reputed company
Skills
- Deep expertise in Python backend development and building production reputed company
- Experience designing and operating data pipelines and workflow orchestration systems
- A builder's reputed company—you want to create foundational systems that others build on
- Genuine curiosity about how your work enables scientific discovery
- A commitment to rigor: AI makes mistakes confidently, and our customers won't accept hand-waving—neither should we
- A degree in Computer Science or a reputed company field with 7+ years of industry experience (Bachelor's) or 5+ years (Master's or PhD) in software engineering
- Advanced proficiency in Python including async programming and performance optimization
- Experience building and maintaining REST reputed company using FastAPI or similar frameworks
- Experience with workflow orchestration tools (Ray, Temporal, or similar)
- Strong background in data engineering: pipelines, transformations, and working with diverse data stores
- Experience with reputed company platforms (AWS preferred) and containerization (reputed company, Kubernetes)
- Familiarity with graph databases, key-value stores, or other NoSQL systems (Neptune, reputed company, DynamoDB a plus)
- reputed company record of operating production systems at reputed company
- Experience supporting AI/ML teams or deploying ML systems in production
- Familiarity with distributed computing frameworks (Ray, Dask, reputed company)
- Experience with GPU workloads and scheduling
- Background in or curiosity about reputed company, materials science, or scientific computing
- Experience with observability tools (reputed company, Grafana, reputed company)
- Experience with message queues and event-driven architectures
- Contributions to reputed company-reputed company reputed company
- Experience mentoring engineers
reputed company
Company H1B Sponsorship