Genesis logo

Data Infrastructure

Genesis

San Francisco, CAJobNo compensation foundPosted 3mo agoVerified open 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
San Francisco, CA
Work Authorization
Not specified

Job overview

The role designs, builds, and maintains large‑scale data pipelines and core data infrastructure for robotics foundation model training, standardizing data models and unifying processing across real‑world and synthetic datasets while collaborating with a driven AI team.

Skills & qualifications

RequiredNice to have

Skills

Data PipelineApache KafkaComputer Data StorageData LakesData WarehousingKubernetesPythonGoSparkKafkaData WarehousesObject Stores (S3)TerraformLarge‑Scale Data PipelinesDistributed Systems

Qualifications

8+ Years Experience

Full job description

What You’ll Do

  • Design, build, and maintain large-scale data pipelines (batch and streaming) for robotics foundation model training and evaluation at petabyte scale

  • Own core data infrastructure: data model, storage systems, ingestion pipelines, transformation frameworks, and orchestration layers

  • Standardize data models and unify processing pipelines across real-world teleoperation and synthetic simulation datasets

  • Collaborate with a team of driven individuals committed to building general-purpose Physical AI

What You’ll Bring

  • Excellent software engineering skills (Python, Go, or similar)

  • Extensive experience designing, building, and maintaining large-scale data pipelines (8+ years)

  • Deep understanding of distributed systems (Spark, Kafka, or similar)

  • Extensive experience with data storage technologies (data lakes, warehouses, object stores like S3)

  • Experience running and maintaining production-grade infrastructure (Kubernetes, Terraform)

  • Bonus: Experience supporting AI systems, in particular embodied AI like self-driving

You've read the whole posting — now see how you match it.