
Principal Software Engineer, Data Processing OSS
San Jose, CA · HybridJob$228–339K/yrSeen 1 day agoSeen in employer's feed 1 day ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Jose, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
The Senior/Principal Data Processing Engineer - OSS will advance open-source platforms and integrate them with DataPelago’s data processing engine. The role focuses on functional breadth, performance, scale, and reliability through upstream and downstream contributions, with opportunities to engage open-source communities. The engineer will work across architecture, core development, performance optimization, and collaboration with engineering, product management, and customer success teams.
Skills & qualifications
Skills
Qualifications
Benefits
Full job description
Job Summary
DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing.
DataPelago Nucleus is a universal data processing engine that is designed to accelerate the processing of diverse data – structured through unstructured – with any parallel processing framework – e.g., Apache Spark – on any infrastructure – including vectorized CPU and GPU. DataPelago Accelerator for Spark (DPA-S), our product based on DataPelago Nucleus, is running large-scale production applications of many globally renowned customers.
The Opportunity:
As a Senior/Principal Data Processing Engineer - OSS, you will be a key individual contributor in adopting and advancing capabilities of open-source software (OSS) platforms such as Apache Gluten, Velox, Apache Spark, and Apache Flink in the context of DataPelago’s data processing engine. You will enhance functional breadth, performance, scale, and reliability of the DataPelago engine through downstream and upstream contributions. You will have the opportunity to engage with community working on these platforms. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers.
What You'll Do:
-
Architect: Influence the architecture of how our data processing engine interfaces with open-source platforms and engines.
-
Design: Lead design of functional and performance enhancements to open-source platforms such as Apache Gluten and Velox and their integration with our data processing engine.
-
Core Development: Individually design, implement, test, optimize, and maintain components of the data processing engine.
-
Innovation and Differentiation: Analyze technology roadmap of Apache Gluten, Velox, and equivalent platforms and identify opportunities for our engine to enhance technology and product leadership.
-
Collaboration: Partner effectively with engineering, product management, open-source community and customer success teams.
-
Continuous Improvement: Foster best practices in design and code reviews, testing, CI/CD, and issue resolution to maintain highest product quality, security, efficiency, & productivity.
What You'll Bring
-
15+ years of relevant experience, Bachelor's degree in Computer Science, or a related field OR a Master's degree in Computer Science or a related field.
-
3+ years of deep technical experience in developing core components of Apache Spark, Apache Flink, Apache Doris, Apache Gluten, Velox, Apache DataFusion, Apache DataFusion Comet, or equivalent platforms designed for large-scale data processing.
-
3+ years of deep technical experience in instrumenting, analyzing, and optimizing the performance of data processing engine components on benchmark and customer workloads.
-
Sound knowledge of the architecture and internal operation of one or more of Apache Spark, Apache Flink, Presto/Trino.
-
Demonstrated experience in the design, development, and successful release of high-performance data processing engines for large production deployments.
-
Exceptional programming skills in C, C++, and Java.
-
Extensive development experience in Linux environments.
-
Strong analytical and problem-solving skills with a passion for performance optimization.
Compensation:
The target salary range for this position is 227,800 - 338,800 USD. The salary offered will be determined by the candidate's location, qualifications, experience, and education and may be outside of this range. Final compensation packages are competitive and in line with industry standards, reflecting a variety of factors, and include a comprehensive benefits package. This may cover Health Insurance, Life Insurance, Retirement or Pension Plans, Paid Time Off, various Leave options, Performance-Based Incentives, employee stock purchase plan, and/or restricted stocks (RSU’s), with all offerings subject to regional variations and governed by local laws, regulations, and company policies. Benefits may vary by country and region, and further details will be provided as part of the recruitment process.
136539
We are all about helping customers turn challenges into business opportunity. It starts with bringing new thinking to age-old problems, like how to use data most effectively to run better - but also to innovate. We tailor our approach to the customer's unique needs with a combination of fresh thinking and proven approaches.
At NetApp, we embrace a hybrid working environment designed to strengthen connection, collaboration, and culture for all employees. This means that most roles will have some level of in-office and/or in-person expectations, which will be shared during the recruitment process.
Equal Opportunity Employer:
NetApp is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination based on age, race, color, gender, sexual orientation, gender identity, national origin, religion, disability or genetic information, pregnancy, protected veteran status, and any other protected classification.
Why You'll Thrive at NetApp
At NetApp, you won't wait for the perfect moment—you'll make it. The early planning, the extra thought, the bold idea that turns good into great: That's how our people operate and how we continue to push the boundaries of data infrastructure.
NetApp is the trusted partner for organizations transforming data into opportunity. As the only enterprise-grade storage service natively embedded in Google Cloud, AWS, and Microsoft Azure, we empower customers to run everything from traditional workloads to enterprise AI with unmatched performance, resilience, and security.
Our culture
We celebrate mold breakers, bold thinkers, and problem solvers. We reward initiative, impact, and ownership. We provide flexibility so you can balance professional ambition with your personal life. Here, differences are not just welcomed—they drive everything we do.
If you're ready to innovate, rise to the challenge, and own every moment - make your next move your best one. Apply now.
Submitting an Application
To ensure a streamlined and fair hiring process for all candidates, our team only reviews applications submitted through our company website. This practice allows us to track, assess, and respond to applicants efficiently. Emailing our employees, recruiters, or Human Resources personnel directly will not influence your application.
AI Disclosure
For select roles, some stages of our hiring process may use artificial intelligence tools to help evaluate applications and candidate selection. These tools support—rather than replace—human decision-making.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Principal EngineerFujifilm · Santa Clara, CA · $160–200K/yrPosted 3w agoPosted 3w ago
Principal Software Development Engineer (Microservices)Zscaler · Santa Clara, CA (Hybrid) · $186–265K/yrPosted 2 days agoPosted 2 days ago
Principal Software Engineer 2 - OLTPSnowflake Inc. · Menlo Park, CA (Hybrid) · $304–380K/yrPosted 3 days agoPosted 3 days ago- Lead Architect / Principal EngineerLitmus · Santa Clara, CA (Hybrid)Posted 1w agoPosted 1w ago
Principal Software Engineer – 5G Core NetworksRTX Corporation · san jose, CA · $118–225K/yrPosted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.