Member of Technical Staff - ML Performance
San Francisco, CAFull-timePosted 1y agoStill listed 3 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Modal Labs is hiring a Member of Technical Staff - ML Performance. Modal is building a new infrastructure layer for AI, supporting category-defining companies with instant GPU access, sub-second container starts, and native storage. The company has raised $355M Series C at a $4.65B valuation, crossed $300M+ ARR, and grown fivefold since September. The team includes creators of popular open-source projects, academic researchers, and experienced engineering and product leaders. They are seeking strong engineers to make ML systems performant at scale.
Successful candidates bring 5+ Years High-Performance Coding Experience. Important skills include PyTorch, High-Level ML Frameworks, vLLM, TensorRT, Nvidia GPU Architecture, and CUDA. Preferred (not required): Low-Level Operating System Foundations, Linux Kernel, File Systems, and Containers.
Skills & qualifications
Skills
Qualifications
Full job description
About Us: AI needs a new infrastructure layer. We're building it at Modal.
Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.
Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.
Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.
The Role: We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!
Requirements:
-
5+ years of experience writing high-quality, high-performance code.
-
Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
-
Familiarity with Nvidia GPU architecture and CUDA.
-
Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).
-
Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Member of Technical Staff - Inference RuntimeModal · Remote · US · $220–300K/yrPosted 4 days agoPosted 4 days agoMember of Technical StaffLabelbox · San Francisco, CA · $140–200K/yrPosted 2w agoPosted 2w ago
Member of Technical Staff — Machine Learning & Agent Security EngineeringSalesforce · Bellevue, WA · $117–177K/yrPosted 2w agoPosted 2w ago
Member of Technical Staff - StorageModal · San Francisco, CA · $250–300K/yrPosted 2w agoPosted 2w ago
Member of Technical StaffCapy · San Francisco, CA · $200–300K/yrPosted 3 days agoPosted 3 days ago
You've read the whole posting — now see how you match it.