RoboForce logo

Senior / Staff AI Research Engineer, Data Infrastructure

RoboForce

CA · HybridFull-time$185–325K/yrPosted 9mo agoStill listed 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$185–325K/yr
Location
CAHybrid
Schedule
Full-time
Work Authorization
Visa required • Visa sponsorship

Olive lists jobs from US employers, including remote roles you can work from the United States.

Requirements

Credentials this posting asks for.

Bachelor's degree

Job overview

RoboForce is hiring a Senior / Staff AI Research Engineer, Data Infrastructure. RoboForce, an AI robotics company focused on industrial deployment, seeks a Senior/Staff AI Research Engineer to own the full data and learning pipeline—from raw teleoperation collection through curation, annotation, storage, and post‑training infrastructure that scores demonstrations and feeds failure data back into model retraining.

Key focus areas include Design and maintain end‑to‑end data collection pipelines ingesting multimodal demonstration data, Build annotation tooling and data curation workflows for quality filtering, deduplication, and episode scoring, and Develop post‑SFT reinforcement learning infrastructure to implement reward scoring and mine failure patterns.

Successful candidates bring Bachelor's Or Master's Degree In Computer Science, Robotics, Or Related Field, 5+ Years Of Experience, and 5 Days/Week In-Office Collaboration. Important skills include Data Collection Pipelines, Multimodal Demonstration Data Ingestion, Data Synchronization, Data Versioning, Distributed Storage, and Annotation Tooling. Preferred (not required): Robotics Data Collection Hardware, Teleoperation Devices, UMI, and GELLO.

Skills & qualifications

RequiredNice to have

Skills

Data Collection PipelinesMultimodal Demonstration Data IngestionData SynchronizationData VersioningDistributed StorageAnnotation ToolingData Curation WorkflowsQuality FilteringDeduplicationEpisode ScoringDomain ReweightingTraining Datasets for Robot Policy LearningPost-SFT Reinforcement Learning InfrastructureReward ScoringFailure Pattern MiningFailure Data CategorizationRetraining Loop IntegrationEvaluation InfrastructureTest InfrastructurePolicy Rollout LoggingStructured Results CaptureActionable Diagnostics SurfacingData Schemas DefinitionEpisode Formats DefinitionPipeline Interfaces DefinitionRapid Iteration SupportVLA TrainingManipulation Policy TrainingScalable Storage SystemsRetrieval SystemsHeterogeneous Robot Data ManagementPythonProduction-Grade Data PipelinesETL SystemsLarge-Scale Dataset ManagementAmazon S3Cloud StorageHDF5WebDatasetZarrPost-Training InfrastructureSFT PipelinesReward ModelingRL Training LoopsPPODPORejection SamplingPyTorchJAXML Training WorkflowsRobotics Data Collection HardwareTeleoperation DevicesUMIGELLORobot Learning PipelinesImitation LearningBehavior CloningVLA/VLM Fine-Tuning WorkflowsExperiment Tracking InfrastructureWeights & BiasesMLflowCustom Rollout LoggersAnnotation Tooling DesignHuman-in-the-Loop Labeling SystemsData PipelineReinforcement Learning

Qualifications

Bachelor's Degree in Computer Science or RoboticsMaster's Degree in Computer Science or Robotics5+ Years Experience5 Days/Week in-Office Collaboration

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match

Full job description

Why RoboForce

RoboForce is an AI robotics company developing Physical AI–powered Robo-Labor for dull, dirty, and dangerous work. The company's robots are engineered for demanding industrial environments, with a focus on real-world deployment and scalability.

We are looking for a Senior / Staff AI Research Engineer, Data Infrastructure to build the data and learning engine behind RoboForce's Physical AI stack. In this role, you will own the full pipeline — from raw teleoperation and UMI device data collection through curation, annotation, and storage, to post-training infrastructure that scores demonstrations, identifies failure patterns, and closes the loop back into model retraining.

Responsibilities

Design and maintain end-to-end data collection pipelines ingesting multimodal demonstration data from teleoperation devices and UMI hardware, including synchronization, versioning, and distributed storage at scale.

Build annotation tooling and data curation workflows — quality filtering, deduplication, episode scoring, and domain reweighting — to produce high-quality training datasets for robot policy learning.

Develop post-SFT reinforcement learning infrastructure: implement reward scoring on demonstrations, mine and categorize failure patterns, and feed curated failure data back into the retraining loop.

Build evaluation and test infrastructure to log policy rollouts on-robot, capture structured results, and surface actionable diagnostics for the research team.

Collaborate with ML researchers to define data schemas, episode formats, and pipeline interfaces that support rapid iteration on VLA and manipulation policy training.

Architect scalable storage and retrieval systems for heterogeneous robot data (vision, proprioception, action, language) across both cloud and on-prem environments.

Requirements

Bachelor's or Master's degree in Computer Science, Robotics, or related field with 5+ years of experience.

Strong proficiency in Python and experience building production-grade data pipelines and ETL systems.

Hands-on experience with large-scale dataset management, including versioning, deduplication, quality filtering, and distributed storage (e.g., S3, GCS, HDF5, WebDataset, Zarr).

Experience building or working with post-training infrastructure — SFT pipelines, reward modeling, or RL training loops (e.g., PPO, DPO, rejection sampling).

Familiarity with deep learning frameworks (PyTorch, JAX) and ML training workflows sufficient to collaborate tightly with research teams.

Requires 5 days/week in-office collaboration with the teams.

Bonus Qualifications

Experience with robotics data collection hardware — teleoperation devices, UMI, GELLO, or similar — and the synchronization and preprocessing challenges they introduce.

Familiarity with robot learning pipelines: imitation learning, behavior cloning, or VLA/VLM fine-tuning workflows.

Experience building evaluation or experiment tracking infrastructure (e.g., Weights & Biases, MLflow, custom rollout loggers).

Proven ability to design annotation tooling or human-in-the-loop labeling systems for structured or multimodal data.

Benefits

Competitive stock options/equity programs.

Health, dental, and vision insurance, 401(k) plan.

Visa sponsorship and green card support for qualified candidates.

Lunches and dinners, a fully stocked kitchen, and regular team-building events.

Compensation: Salary $185,000–$325,000 USD + Bonus + Equity

The base salary range above represents the expected compensation for this full-time U.S. position. Final compensation will be determined based on role scope, level, location, job-related skills, experience, and relevant education or training, and may fall outside the listed range in exceptional cases.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.