
ML Research Engineer, Training
San Francisco, CAFull-timeSeen todayStill listed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Weave Robotics seeks a Machine Learning Research Engineer to build end‑to‑end training stacks, handle terabyte‑scale multimodal robot data, and turn research prototypes into production infrastructure, enabling robots to perform useful work in real homes and businesses.
Skills & qualifications
Skills
Full job description
Join Us, and Ship Robots Weave was founded to build the robots we’d want to have in our own home. We believe the next generation of robotics will transform everyday life by enabling people to do more and to reclaim time to spend on what’s important. We also believe robots are in a sense like any other product: to matter, they have to ship. Our robots are already operating in real homes and businesses, giving us the opportunity to rapidly improve from real-world experience. With a growing team, strong customer demand, and capital for expansion, we’re entering an exciting stage of growth—and we’re looking for people with exceptional talent and standards to help bring home robotics to millions of households. The Role Most robot learning research is graded on evals that don't survive contact with the field. Ours is graded by robots doing useful work in real homes and businesses, every day. We're one of the first companies with a deployed fleet generating real-world robot data at terabyte scale. The pipeline and training stack you build are what turns that data into capability. Model quality is set as much by training as by architecture: what data gets in, how it's sampled, whether the run is stable, or whether a silent bug ate the gradient three days ago. You'll own that layer from raw fleet uploads to the batch that hits the GPU. When the stack is right, ideas become models in training in days, and deployed in weeks. Responsibilities
-
Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.
-
Large-scale data handling: Develop high-throughput data ingestion, transformation, and storage systems capable of processing terabytes of multimodal robot data, including video, proprioception, and sensor streams.
-
Optimize research productivity: Grow the codebase that makes experiments reproducible, scalable, and easy to launch, monitor, retry/recover, and debug.
-
Sampling and curation: Drive sampling and curation decisions that show up in model behavior.
-
Make runs fast and honest: profile and fix throughput bottlenecks, chase down loss spikes and silent data bugs, keep results reproducible enough to trust: from data loading to GPU kernels.
-
From research to production: Turn research prototypes into infrastructure the whole team trains on.
What You'll Bring
-
ML system expertise: Deep PyTorch or JAX experience, including multi-node distributed training (FSDP, DDP, or equivalent) on real workloads.
-
Performance engineering: Experience profiling and optimizing GPU utilization, data pipelines, I/O bottlenecks, memory usage, and distributed training performance, including CUDA-level profiling tools (e.g. Nsight Systems) and NCCL tuning.
-
Training run judgement: You can read a loss curve, tell instability from a data bug and know when to kill a run.
Nice to Have
-
Robot learning exposure: you’ve trained policies (VLAs, world models, RL) and can tell a data problem from a model problem.
-
Cluster and cloud infrastructure experience: Kubernetes, SLURM, GCP/AWS.
-
Large-scale post-training experience: SFT, reward modeling, RL fine-tuning.
-
On-robot inference optimization experience: TensorRT, quantization, distillation.
-
CUDA or Triton kernel work.
-
Experience with video-heavy datasets: transcoding, chunking, and the storage/compute tradeoffs of training on video at scale.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Research Engineer, AI for Chip DesignOpenAI · San Francisco, CA (Hybrid) · $380–500K/yrPosted 2w agoPosted 2w ago
Research Engineer, SafetyDecagon · San Francisco, CA · $200–400K/yrPosted 3w agoPosted 3w ago
Research FellowAbundant · San Francisco, CA · $14,500/moPosted 2w agoPosted 2w ago
Member of Technical Staff, Research (Intern)Abundant · San Francisco, CA · $10,000/moPosted 3w agoPosted 3w agoResearch Engineer - ML InfrastructureChai Discovery · San Francisco, CAPosted 6 days agoPosted 6 days ago
You've read the whole posting — now see how you match it.