Engineer, Supercomputing & Distributed Systems
San Francisco, CAFull-time$95–165K/yrPosted 5mo agoStill listed 3w ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
krea.ai is hiring an Engineer, Supercomputing & Distributed Systems. Krea builds next‑generation AI creative tools, focusing on intuitive, controllable systems for text, images, video, sound, and 3D. The team blends engineers, designers, and artists, operating from a waterfront office in San Francisco, and has raised over $83M to advance its foundation model and infrastructure.
Key focus areas include Design multi‑stage pipelines that turn petabytes of raw data into clean, annotated datasets, Run classification models on billions of images, and Deploy and combine LLMs to caption massive multimedia data.
Important skills include Python, Kubernetes, PyTorch, DuckDB, Arrow, and Distributed Systems Intuition. Preferred (not required): PyArrow, SQL, Massive Relational Databases, and Pandas.
Skills & qualifications
Skills
Benefits
Full job description
About Krea At Krea, we are building next-generation AI creative tools.
We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2 , our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.
We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.
Supercomputing / AI Infra at Krea We build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.
Example projects: Distributed data systems
-
Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
-
Run classification models on billions of images
-
Deploy and combine LLMs to caption massive multimedia data
GPU infrastructure
-
Manage distributed training and inference on 1000+ GPU Kubernetes clusters
-
Solve orchestration and scaling for large-scale GPU job processing
-
Scale workloads and research between clusters in multiple datacenters
Distributed training
-
Profile and optimize dataloaders streaming thousands of images per second
-
Profile and debug InfiniBand networking on huge training runs
-
Build fault tolerance systems for large-scale pretraining
-
Collaborate with researchers on evolving RL infrastructure
Applied ML pipelines
-
Find clean scenes in millions of videos using distributed shot-boundary detection
-
Customize and train models to filter billions of images for questions like "is this a screenshot?"
-
Build the systems that bridge raw cluster capacity and research output
Who we're looking for Systems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought "I want to do that" — this is that team.
You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.
Strong candidates may have experience with…
-
Python, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy…
-
Kubernetes
-
Designing and implementing large-scale ETL systems
-
Fundamental knowledge of containerization, operating systems, file-systems, and networking
-
Distributed systems design
-
Distributed training systems (NCCL, InfiniBand, RDMA)
-
Streaming and event processing systems (Kafka, Pulsar, or similar)
-
PyTorch internals, custom dataloaders, and training infrastructure
What we offer
-
Team : Work alongside a world-class team building the future of AI creative tooling
-
Impact : Significant scope and company-wide impact
-
Competitive compensation : generous salary & equity packages
-
Health & wellness : 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage
-
Time off : Flexible PTO policy
-
Financial planning : 401k with a 4% company-sponsored match
-
Meals in the office : breakfast, lunch, dinner - you name it, we'll cover it
-
Transit : Ubers covered to & from the office
-
Sponsorship : We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).
-
And more!
Please note the above benefits & perks are for full-time employees
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Software Engineer- Infrastructure & Distributed SystemsSalesforce · San Francisco, CA · $149–224K/yrPosted 5 days agoPosted 5 days ago
Lead Software Engineer, Distributed File Systems/Cloud PlatformSalesforce · San Francisco, CA · $208–286K/yrPosted 5 days agoPosted 5 days ago
Software Engineering PMTS, Slack Distributed Data ServiceSalesforce · San Francisco, CA · $197–314K/yrPosted 6 days agoPosted 6 days agoSoftware Engineer, Model RuntimeOpenAI · San Francisco, CA (Hybrid) · $266–445K/yrPosted 2w agoPosted 2w ago
Staff Engineer, Search SystemsMongoDB · San Francisco, CA (Hybrid) · $151–297K/yrPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.