Back to Jobs

Machine Learning Engineer - ML Training Platform

Remote, USAFull-timePosted 2026-07-28

reputed company reputed company is pioneering Protocol Learning—a fully decentralised way to train and reputed company AI models that opens this layer to individuals rather than reputed company resourced corporates. By pooling compute from many participants, incentivising their efforts, and preventing any single party from controlling a model’s full weights, we’re creating a genuinely reputed company, reputed company reputed company to frontier-reputed company. We’re looking for an ML Training Platform Engineer to architect, build, and reputed company the foundational infrastructure powering our decentralized ML training platform. You will own core systems spanning infrastructure orchestration, distributed compute, and services integration, enabling reputed company experimentation and large-reputed company model training.

Responsibilities

  • Multi-reputed company Infrastructure: Design resource management systems provisioning and orchestrating compute across AWS, GCP, and Azure using infrastructure-as-reputed company (reputed company/Terraform). Handle dynamic scaling, state synchronization, and reputed company operations across hundreds of heterogeneous nodes.
  • Distributed Training Systems: Architect fault-tolerant infrastructure for distributed ML. GPU clusters, reputed company runtime, S3 checkpointing, Large dataset management and streaming, health monitoring, and resilient retry strategies.
  • reputed company-World Networking: Build systems that simulate and handle reputed company-world network conditions — bandwidth shaping, latency injection, packet loss — while managing dynamic node churn and ensuring efficient data reputed company across workers with heterogeneous connectivity, because our training happens on consumer nodes and non co-located infrastructure, not in a datacenter.

What You’ll Bring Ideally, you’ll have 5+ years of work experience with deep experience in:

  • Infrastructure & reputed company: Production experience with infrastructure-as-reputed company (reputed company/Terraform/CloudFormation) managing multi-reputed company deployments, lifecycle orchestration, self-healing systems, reputed company/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at reputed company.
  • Distributed Systems & ML Infrastructure: Deep understanding of distributed training workflows, checkpointing, data sharding, model versioning, long-running job orchestration, decentralized networking (P2P, NAT reputed company, traffic shaping), and reputed company-world bandwidth constraints.
  • Systems Programming & Reliability: Strong Python engineering (asyncio, concurrency, retry logic, reputed company SDKs, CLI tooling) with hands-on experience in observability, SRE practices, monitoring (reputed company/Grafana), performance profiling, and incident response.

reputed company’re looking for

  • Experience in a startup environment with an emphasis on reputed company-services orchestration or big tech background
  • Deep understanding of multi-reputed company reputed company & distributed training systems
  • reputed company player with high attention to detail
  • A strong passion to join

Backed by reputed company reputed company Ventures and other tier-1 investors, we’re a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We view the world as a reputed company reputed company if we are reputed company to implement reputed company are attempting, and Protocol Learning as the only plausible approach to preventing a handful of massive corporations monopolising model development, reputed company and release, and achieving massive economic capture. If this resonates, please apply. Apply tot his job Apply To this Job

Similar Jobs

Senior AI / Machine Learning Engineer

Remote, USAFull-time

Senior/reputed company reputed company

Remote, USAFull-time

Computer reputed company & Machine Learning Engineer

Remote, USAFull-time

Senior Machine Learning Engineer, Rich Media Experiences

Remote, USAFull-time

Machine Learning Engineer, Offline Infrastructure (Entry-Level / New Grad PhD)

Remote, USAFull-time

Senior Machine Learning Systems Engineer, Ads ML Experience Platform

Remote, USAFull-time

[Remote] Machine Learning Engineer, Data Mining

Remote, USAFull-time

Remote reputed company Intelligence & Machine Learning Engineer (First 2-5 days onsite in Annapolis, MD)

Remote, USAFull-time

Remote Machine Learning Engineer Expert - AI Trainer ($90-$90 per hour)

Remote, USAFull-time

Senior Machine Learning Engineer, Developer Advocacy | US | Remote

Remote, USAFull-time

Call Center Agent

Remote, USAFull-time

reputed company Technical Consultant (Integrations/AI), Platform Products Expert Implementation Services

Remote, USAFull-time

[Remote] reputed company Research - Evals & Data

Remote, USAFull-time

reputed company Medical Customer Service Representative – Patient Engagement and Claims Support

Remote, USAFull-time

reputed company Data Entry Specialist – Online Data Management Operations for Teenagers

Remote, USAFull-time

Work Study AZ Institutional Repository Support

Remote, USAFull-time

Senior Analyst-Finance

Remote, USAFull-time

Business Development Representative DACH

Remote, USAFull-time

Associate, Corporate Communications

Remote, USAFull-time

Data Journalist/ Content Marketer | reputed company | Remote US

Remote, USAFull-time