Principal AI Research Engineer - RL
San Francisco, CAFull-timePosted 1y agoStill listed 4 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Reflex Robotics is hiring a Principal AI Research Engineer - RL. Reflex Robotics is a three-year-old startup building affordable wheeled humanoid robots to automate dangerous and repetitive tasks in manufacturing and logistics. The company envisions a future where intelligent robots handle boring work. They are backed by Khosla Ventures and have significant revenue lined up pending successful pilots in 2025. Reflex Robotics is looking for stellar on-policy RL engineers to create robust robot policies.
Key focus areas include Re‑implement core RL algorithms such as SAC and DDPG, Debug unstable gradients and tune hyperparameters, and Contribute to sample‑efficient RL research like DreamerV3 and MuZero.
Successful candidates bring Re-Implemented Core RL Algorithms From Scratch, Meaningful Intellectual Contributions To Sample-Efficient RL Algorithms, and Shipped On-Policy RL On Hardware That Learns In The Real-World. Important skills include SAC, DDPG, Debug Unstable Gradients, Tune Hyperparameters, DreamerV3, and MuZero. Preferred (not required): Insane Work Ethic, Push Past Status Quo, Passion For Work, and Desire To Win.
Skills & qualifications
Skills
Qualifications
Full job description
Company Overview
Reflex Robotics is building affordable ($10k) wheeled humanoid robots to automate dangerous and repetitive tasks in manufacturing and logistics.
We envision a future where intelligent robots are doing all kinds of boring work that people hate doing-loading chicken nuggets into Costco boxes, lifting forty pound bags of dog food at Petco stores, and cleaning up cranberry juice spills in your apartment.
We are a four-year-old startup backed by Khosla Ventures, with $100M+/year of revenue lined up pending successful pilots with e-commerce warehouses.
How Does It Work?
Our robots are designed and built entirely in-house by an engineering team that led development of the Stretch robot at Boston Dynamics and key systems on the Tesla Model S, X, and Y production lines. Reflex robots are high-performance, low-inertia, and optimized for low-cost manufacturing.
We’ve built the best real-time teleoperation system in the world, allowing a remote operator in South America to “play a video game” to control our robots at human-level speeds . This has allowed us to already ship robots with positive unit economics, and enables us to create a powerful human-intervention + RL product feedback loop.
Our system allows us to collect high-quality demonstrations at scale—giving us the proprietary data engine needed to train increasingly capable AI systems. We're on track to build the largest robotics dataset in the world, which will serve as an important long-term advantage.
Key Company Beliefs
-
High-quality, proprietary robotics data is the next foundation for generational AI companies (like Tesla FSD and ChatGPT).
-
Being nerd-sniped by maximizing an engineering metric is way less important than solving our customers’ biggest pain points.
-
An insane work ethic is required for outsized success—and you'll be rewarded for it.
What We’re Looking For
We’re looking for stellar on-policy RL engineers to work on creating robust robot policies.
We’re still a small team—which means high ownership, high equity, and the chance to shape the product from the ground up.
VLAs and other great “base policies” for robotics achieve ~80% success rates, but in real robot deployments, it’s essential to achieve 99.99% success rates. We can’t ask our customers to tolerate our robots packing three socks into a bin instead of four, or swapping shipping labels between two packages—not even once!
You should apply for this role if:
-
You’ve re-implemented core RL algorithms (SAC, DDPG) from scratch and can debug unstable gradients / tune hyperparameters correctly
-
You’ve made meaningful intellectual contributions to sample-efficient RL algorithms (e.g., DreamerV3 and MuZero)
-
You’ve shipped on-policy RL on hardware that learns in the real-world (e.g., for quadruped walking or drone racing)
You’d be joining a company that already has a solid core business—with working hardware, delighted customers, and profitable unit economics. Reflex is de-risked enough to see the hazy outlines of success, but still small enough that there’s enormous upside up for grabs.
Come Join Us
This is a rare opportunity to help build a flagship robotics company from the ground up—and to do work that will truly matter, reshaping what people believe is possible in robotics.
We love to see the things you’ve worked on. Have a portfolio or insane project you’ve worked on? Share it. We’re looking for people who push past the status quo, are passionate at work and in their own time—we’re looking for people who want to win.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Principal Engineer, AI/ML SecurityDigitalOcean · Remote · US · $230–288K/yrPosted 2w agoPosted 2w ago
SVP, Research & Insights - Experience DesignSalesforce · San Francisco, CA · $399–453K/yrPosted 2w agoPosted 2w agoResearch Engineer, AI for Chip DesignOpenAI · San Francisco, CA (Hybrid) · $380–500K/yrPosted 2w agoPosted 2w ago
AI/ML Research InternAfterQuery · San Francisco, CA · $8,000/moPosted 2w agoPosted 2w agoResearch Program Manager, Frontier AssuranceOpenAI · San Francisco, CA (Hybrid) · $239–328K/yrPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.