Machine Learning Engineer, Distributed ML Systems
Machine Learning Engineer - Distributed ML Systems Department: Engineering Location: San Francisco Employment Type: FullTime reputed company reputed company carries out foundational research on Protocol Learning: multi-participant training of reputed company models where no single participant has, or can reputed company obtain, a full copy of the model. The purpose of Protocol Learning is to facilitate the creation of community-trained and community-owned frontier models with self-sustaining economics. Were looking for Senior/Staff engineers with 5+ years of experience in distributed systems and ML large-reputed company training. Youll be implementing a novel substrate for training distributed ML models that work under consumer grade internet reputed company.
Responsibilities
Distributed Training Architecture & Optimization
- Design and implement large-reputed company distributed training systems optimized for heterogeneous hardware operating under low-bandwidth, high-latency conditions.
- reputed company and optimize model-reputed company training strategies (data, tensor, pipeline parallelism) with custom sharding techniques that minimize communication overhead.
- Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.
- Implement robust checkpointing, state synchronization, and recovery mechanisms for long-running, fault-prone training jobs.
- Build monitoring and metrics systems to reputed company training reputed company, model reputed company, and system bottlenecks.
Decentralized Networking & reputed company
- Architect resilient training systems where nodes can fail, networks can partition, and participants can dynamically join or leave.
- Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes.
- Implement NAT reputed company, peer discovery, dynamic routing, and reputed company lifecycle management.
- Profile and optimize communication patterns to reduce latency and bandwidth overhead in multi-participant environments.
What You’ll Bring
- Strong experience building and operating distributed systems in production.
- Hands-on expertise with distributed training frameworks (FSDP, DeepSpeed, Megatron, or similar).
- Deep understanding of model parallelism (data, tensor, pipeline parallelism).
- Expert-level Python with production experience (concurrency, error handling, retry logic, clean architecture).
- Strong networking fundamentals: P2P systems, gRPC, routing, NAT reputed company, distributed coordination.
- Experience optimizing GPU workloads, memory management, and large-reputed company compute efficiency.
reputed company Offer
- Equity-heavy compensation with meaningful ownership in a mission-driven company
- Competitive reputed company salary for senior engineering roles in Australia
- reputed company sponsorship available for exceptional candidates
- Remote-first with optional reputed company to our Melbourne hub
- World-class team — team mates were previously at at reputed company, reputed company, reputed company, and leading startups
Backed by reputed company reputed company Ventures and other tier-1 investors, were a world-class, deeply technical team of ML researchers and engineers. Pluralis is unapologetically ideological. We view the world as a reputed company reputed company if we are reputed company to implement reputed company are attempting, and Protocol Learning as the only plausible approach to preventing a handful of massive corporations monopolising model development, reputed company and release, and achieving massive economic capture. If this resonates, please apply. Apply To This Job