ML Infrastructure Engineer
Redwood City, CAJobPosted 8mo agoStill listed 4 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Redwood City, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Sunday is building personal home robots and seeks passionate engineers to develop end-to-end ML models and foundational systems that accelerate robot deployment. The role offers a chance to shape data pipelines, training infrastructure, and real-time inference, working closely with researchers to bring generalized robots into households.
Skills & qualifications
Skills
Full job description
Join Us in Building the Future of Home Robotics At Sunday, we're developing personal robots to reclaim the hours lost to repetitive tasks. We're focused on an ambitious goal to make generalized robots broadly accessible, enabling households to take back quality time. We have spent the last 18 months building a talented team, securing capital, and validating our technology. We are now seeking passionate individuals to join us in the next phase of our growth. If you are ready to apply your skills to the forefront of robotics innovation, we’d love to hear from you.
The Role Sunday Robotics is building the future of home robotics. We're developing end-to-end ML models for robot manipulation, and you'll have the opportunity to build and shape foundational systems that directly accelerate our path to putting robots in homes. This is a broad role that can be tailored to your specific area of expertise: data pipelines, training infrastructure or inference. You'll build systems across the full robot learning pipeline: ingesting and processing multimodal data, scaling distributed training, optimizing inference for real-time control and building research tooling.
What You'll Do Training and Inference Infrastructure
-
Maintain an effective research codebase with good ergonomics, optimizing for fast iteration and correctness
-
Own infrastructure for model training: job scheduling, checkpointing, metrics, and logging
-
Scale distributed training across GPU clusters with minimal researcher friction
-
Enable training of larger models through sharding, activation checkpointing and memory optimization
-
Profile and optimize gpu utilization, memory usage and training throughput
-
Build low-latency inference pipeline for real-time robot control, apply quantization, distillation and model compilation to optimize inference performance
-
Work closely with researchers and roboticists to translate research needs into reliable software and infrastructure
Data Pipelines and Research Tooling
-
Design high-throughput pipelines for ingesting, validating, and transforming multimodal robot data (video, proprioception, actions)
-
Build storage systems and metadata indexing for efficient dataset management at large scale
-
Optimize dataloaders, sharding and prefetching to minimize time from data arrival to model training
-
Build research tooling for debugging, visualization and experiment analysis
What We're Looking For
-
Strong software engineering and systems fundamentals
-
Experience building distributed systems or large-scale data pipelines
-
Hands-on experience with ML training infrastructure, ideally PyTorch
-
Comfort reasoning about performance, memory, I/O, and GPU utilization
-
Experience managing training workloads (SLURM, Kubernetes, or similar)
-
Ownership mindset: you design, build, operate, and iterate on systems end-to-end
-
Enjoy working closely with researchers and unblocking fast-moving projects
Nice to Have
-
Experience with robotics data pipelines or multimodal models
-
Background in VLAs, Video Generation architectures or robot learning systems
-
Deep ML systems experience: training compilers, custom kernels, runtime optimization
-
Hands-on GPU performance tuning
-
Experience with serialization formats for high-performance systems (Protobuf, FlatBuffers, MCAP)
At Sunday Robotics, we’re building technology shaped by real people — curious, creative, and diverse. We’re proud to be an equal opportunity employer and consider all qualified applicants regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status. Even if you don’t meet every single requirement, we encourage you to apply. Studies show that women and underrepresented groups often hold back unless they meet 100% of the criteria — we don’t want that to be the reason we miss out on great talent.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Research Engineer - ML InfrastructureChai Discovery · San Francisco, CAPosted 2w agoPosted 2w ago
ML Infra Engineer, PlatformPhysical Intelligence · San Francisco, CAPosted todayPosted today
Senior Software Engineer, ML Infra - Asset SafetyRoblox · San Mateo, CA · $279–329K/yrPosted 2w agoPosted 2w ago
Sr. Staff Software Engineer, Product ML InfrastructurePinterest · Palo Alto, CA (Hybrid) · $245–429K/yrPosted 1w agoPosted 1w ago
2027 Internship Onboard Infrastructure Engineer, ML InferenceBedrock Robotics · San Francisco, CAPosted 4w agoPosted 4w ago
You've read the whole posting — now see how you match it.