
Software Engineer, Distributed Training
Palo Alto, CAJob$200–420K/yrPosted 1 day agoStill listed 1 day ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Palo Alto, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
River AI seeks exceptional systems engineers to build distributed training engines for the River API, focusing on fine‑tuning and reinforcement learning across large GPU clusters. The role involves owning training workload execution, optimizing memory and parallelism, implementing reliable checkpointing, and collaborating with researchers to bring new algorithms to production.
Skills & qualifications
Skills
Qualifications
Benefits
Full job description
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.
Who we are
We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
About the Role
We are looking for exceptional systems engineers to build the distributed training engines behind the River API. Your goal is to make fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.
You will own the execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Working closely with researchers and inference engineers, you will bring new learning methods into production and improve how efficiently models use compute.
What You’ll Do
-
Build and optimize distributed training for large dense and mixture-of-experts models, including low-rank adapter training.
-
Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
-
Optimize GPU memory use, parallelism, and communication to increase training throughput.
-
Implement reliable checkpointing, resumption, and worker recovery while preserving consistent training state.
-
Validate losses, gradients, and optimizer behavior, and diagnose numerical or distributed execution failures.
-
Partner with researchers to implement new algorithms and make them accessible through the River API.
Skills & Qualifications
Minimum Qualifications:
-
Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
-
Hands-on experience building or substantially improving distributed model-training systems.
-
Strong proficiency in Python and a modern deep-learning framework, such as PyTorch or JAX.
-
Solid understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management.
-
Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery.
-
A highly collaborative mindset and a bias for action to push boundaries across the stack.
Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)
-
Experience with reinforcement-learning infrastructure, rollout generation, or asynchronous training.
-
Familiarity with tensor, pipeline, expert, or data parallelism and their performance tradeoffs.
-
Work on LoRA, mixture-of-experts training, distributed optimizers, or activation checkpointing.
-
Experience with NCCL, communication profiling, and overlapping computation with data transfers.
-
Proficiency in C++, Rust, or CUDA, with experience investigating performance below the framework layer.
-
Contributions to training frameworks or a track record of operating large training runs.
Logistics & Benefits
-
Location: Palo Alto, California.
-
Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.
-
Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
-
Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
- Senior Software EngineerRobinhood · Menlo Park, CA (Hybrid)Posted 3w agoPosted 3w ago
Senior Software Engineer(Distributed Systems)Workday · Pleasanton, CA (Flexible) · $190–285K/yrPosted 1w agoPosted 1w agoCluster Operations Software EngineerCerebras · Sunnyvale, CAPosted 4w agoPosted 4w ago
Principal Software Engineer, Post TrainingWaymo · Mountain View, CA (Hybrid) · $349–431K/yrPosted 2w agoPosted 2w ago
Principal Software EngineerCisco · Milpitas, CA · $251–363K/yrPosted 5 days agoPosted 5 days ago
You've read the whole posting — now see how you match it.