LanceDB logo

Senior Software Engineer

LanceDB

ORRemoteFull-time$180–250K/yrPosted 10mo agoVerified open 6 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$180–250K/yr
Location
ORRemote
Schedule
Full-time
Work Authorization
Not specified

Job overview

LanceDB is hiring a Senior Software Engineer. LanceDB seeks a Senior Software Engineer to expand the reach of its data platform within the broader data infrastructure ecosystem, contributing to high‑performance computing, big data, and open‑source systems while improving scalability, performance, and usability.

Key focus areas include Design and maintain efficient distributed Lance dataset operations, Build efficient indices to enable predicate pushdown and accelerate queries in Spark, Ray, or Trino, and Work on table formats, data encodings, and various aspects of the Lance format in Rust.

Successful candidates bring 10+ Years Of Experience Building High-Performance Databases, Big Data Systems, Or Large-Scale Data Services, Deep Understanding Of Internals Of Open-Source Big Data Or AI Training Systems, and Willingness To Learn Rust. Important skills include Distributed Lance Dataset Operations, Predicate Pushdown, Apache Spark, Ray, Trino, and Table Formats. Preferred (not required): Collaboration, Apache Contributor, Apache Committer, and Apache PMC Member.

Skills & qualifications

RequiredNice to have

Skills

Distributed Lance Dataset OperationsPredicate PushdownApache SparkRayTrinoTable FormatsData EncodingsLance FormatRustOpen-Source Community IntegrationHive MetastorePrestoData Infrastructure SystemsInternal Data Processing InfrastructureOpen-Source PromotionHigh-Performance DatabasesBig Data SystemsLarge-Scale Data ServicesAI Training Systems InternalsHadoopApache FlinkIcebergDelta LakeHudiClickHousePyTorchJAXHigh-Performance ComputingC++JavaScalaMove FastWork IndependentlyCollaborationApache ContributorApache CommitterApache PMC MemberOpen-Source Projects ContributorApache ArrowDataFusionParquetDriving Large FeaturesIntegrations in Distributed SystemsCommunity PresenceBuilding Efficient IndicesOpen-Source Community EffortsOperating Internal Data Processing InfrastructureImproving Internal Data Processing InfrastructureBig Data ConferencesCollaborate With High-Caliber TeamOpen-Source Collaboration

Qualifications

10+ Years of Experience Building High-Performance Databases, Big Data Systems, or Large-Scale Data ServicesDeep Understanding of Internals of Open-Source Big Data or AI Training SystemsWillingness to Learn RustContributor, Committer, or PMC Member in Apache or Other Large Open-Source Projects

Full job description

ABOUT LANCEDB

LanceDB http://lancedb.com/ is the preeminent data platform for multimodal AI use cases. From hyper-scalable vector search to advanced retrieval for RAG, from streaming training data to interactive exploration of large-scale AI datasets, LanceDB is the best foundation for your AI application, and powers some of the most groundbreaking applications and challenging requirements today.

ABOUT THE ROLE

We’re looking for a Senior Software Engineer to help expand the reach of Lance and LanceDB within the broader data infrastructure ecosystem. You’ll work at the intersection of high-performance computing, big data, and open-source systems. You will contribute scale and performance improvements, integrations with the wider data and AI ecosystem, simplifying distributed operations, and usability and maintainability enhancements.

YOU’LL BE RESPONSIBLE FOR

  • Designing and maintaining efficient distributed Lance dataset operations

  • Building efficient indices to enable predicate pushdown and accelerate queries in Spark, Ray, or Trino

  • Working on table formats, data encodings, and various aspects of the Lance format in Rust

  • Driving open-source community efforts to integrate the Lance format with Spark, Hive Metastore, Presto, Trino, Ray, and other data infrastructure systems

  • Operating and improving internal data processing infrastructure

  • Promoting the Lance format in open-source communities and at Big Data conferences

REQUIREMENTS

  • 10+ years of experience building high-performance databases, big data systems, or large-scale data services

  • Deep understanding of internals of open-source Big Data or AI training systems (e.g., Hadoop, Spark, Flink, Ray, Iceberg, Delta Lake, Hudi, ClickHouse, Trino, Presto, PyTorch, or JAX)

  • Strong experience with high-performance computing in C++, Java, and/or Scala

  • Experience with Rust (or willingness to learn it)

  • Proven ability to move fast, work independently, and collaborate with a high-caliber team

NICE TO HAVE

  • Contributor, committer, or PMC member in Apache or other large open-source projects

  • Experience with Apache Arrow, DataFusion, Parquet, Iceberg, or Delta Lake

  • Track record of driving large features or integrations in distributed systems

  • Strong community presence and passion for open-source collaboration

WHAT WE OFFER

  • A key role shaping an open-source project with real production usage

  • Remote-first team with flexible hours

  • Competitive compensation, equity, and benefits

  • Generous learning budget and support for open-source contributions

WHY JOIN US

You’ll join a world-class team of open-source builders, including co-authors of pandas, and contributors to HDFS, Arrow, Iceberg, and HBase. You’ll collaborate on systems that power next-generation AI workloads while shaping how LanceDB operates and scales production environments.

You've read the whole posting — now see how you match it.