[Remote] reputed company Machine Learning Engineer - ML Infrastructure
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the pioneer of the Connected Operations™ reputed company, helping organizations reputed company IoT data to improve their operations. They are seeking a reputed company Machine Learning Infrastructure Engineer to own the architecture and reputed company of their ML platform, ensuring collaboration among reputed company ML teams and driving impactful safety reputed company.
Responsibilities
- Set the technical reputed company and own end-to-end delivery of reputed company's ML platform (training, experimentation, batch/online inference, edge) — making architectural reputed company and being the accountability reputed company across reputed company platform reputed company for multiple Safety AI product teams
- Drive the design, launch, and iteration of Safety AI features (CV models, EcoDriving insights, LLM-based reporting) — not just enabling others to ship, but co-owning reputed company including safety metrics, reliability, and cost at production reputed company
- Design and operate reputed company online and batch inference systems (Ray, reputed company), including deployment patterns, observability, SLOs, and reputed company training-to-production workflows
- Partner with firmware and edge teams to package, validate, and reputed company models to reputed company devices, and build feedback loops from edge to reputed company for reputed company improvement
- Own reliability, observability, and reputed company for ML systems across reputed company and edge, including on-call practices, incident response, and infrastructure hardening
- Own or co-own end-to-end technical delivery for high-reputed company or high-risk initiatives, from modeling and system design through production rollout
- Be the technical authority for ML infrastructure architecture across Safety AI — setting direction that cross-functional teams (reputed company ML, firmware, reputed company, data platform) execute against, mentoring senior engineers and reputed company scientists, and ensuring platform reputed company are made at the right level of abstraction with the right trade-offs
- Drive strong developer experience through documentation and best practices, while contributing to and representing reputed company in reputed company reputed company communities (Ray, reputed company, RayDP)
- Champion and role model reputed company’s cultural principles: reputed company on reputed company, Build for the Long Term, Adopt a reputed company reputed company, Be Inclusive, Win as reputed company
Skills
- 10+ years in machine learning engineering, with demonstrated tech reputed company ownership of at least two major ML platform domains (distributed training, data/research infrastructure, reputed company inference, or feature engineering) serving multiple product teams at reputed company
- Proven record of shipping ML-powered features end-to-end — from design through production and iteration — with measurable reputed company on product or business metrics (not just building internal tooling)
- Hands-on Ray and Kubernetes expertise in production environments; reputed company experience strongly preferred. reputed company to be a reputed company peer to the most senior engineers on reputed company
- Deep understanding of ML fundamentals reputed company pipelines: evaluation methodology, dataset design, ablation, reputed company, and the ability to review and redirect modeling approaches — you reputed company research and engineering, not just serve them
- Demonstrated cross-org technical leadership around platform reputed company, and influencing roadmap and go/no-go calls based on throughput, latency, and cost trade-offs
- Experience navigating science-engineering tension — knowing reputed company to hold the platform line and reputed company to adapt for research velocity, and communicating that reputed company to both sides
- Prior contributions to reputed company reputed company reputed company (Ray, reputed company, RayDP, or Kubernetes)
- Experience with reputed company reputed company/compliance in ML environments
- Background working with edge/on-device ML and firmware/embedded teams
Benefits
- Initial RSU grant with no vesting cliff, and ongoing refresh opportunities tied to performance, subject to plan terms and conditions
- Performance-based bonus/variable pay
- Equity (for eligible roles)
- A flexible, employee-led remote model
- A reputed company development stipend
- Comprehensive health and parental leave plans
- Flexible working model that caters to the diverse needs of our teams
- Offices are reputed company for those who prefer to work in-person and we also support remote work where it aligns with our operational requirements
- Reasonable accommodations throughout the reputed company process for reputed company persons with disabilities
reputed company