Data Infrastructure
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Job overview
The role designs, builds, and maintains large‑scale data pipelines and core data infrastructure for robotics foundation model training, standardizing data models and unifying processing across real‑world and synthetic datasets while collaborating with a driven AI team.
Skills & qualifications
Skills
Qualifications
Full job description
What You’ll Do
-
Design, build, and maintain large-scale data pipelines (batch and streaming) for robotics foundation model training and evaluation at petabyte scale
-
Own core data infrastructure: data model, storage systems, ingestion pipelines, transformation frameworks, and orchestration layers
-
Standardize data models and unify processing pipelines across real-world teleoperation and synthetic simulation datasets
-
Collaborate with a team of driven individuals committed to building general-purpose Physical AI
What You’ll Bring
-
Excellent software engineering skills (Python, Go, or similar)
-
Extensive experience designing, building, and maintaining large-scale data pipelines (8+ years)
-
Deep understanding of distributed systems (Spark, Kafka, or similar)
-
Extensive experience with data storage technologies (data lakes, warehouses, object stores like S3)
-
Experience running and maintaining production-grade infrastructure (Kubernetes, Terraform)
-
Bonus: Experience supporting AI systems, in particular embodied AI like self-driving
You've read the whole posting — now see how you match it.