GPU Performance Engineer
San Francisco, CAFull-time$125–275K/yrPosted 1y agoStill listed 1 day ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
Genmo is hiring a GPU Performance Engineer. Genmo, a research lab focused on open video generation models, seeks a GPU Performance Engineer to maximize H100 infrastructure performance, optimize model serving, and achieve significant speedups through advanced profiling and custom kernel development.
Key focus areas include Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation, Write high-performance CUDA and Triton kernels for critical model operations, and Optimize cold start latency from seconds to milliseconds for serving infrastructure.
Successful candidates bring Bachelor's Or Master's Degree In Computer Science, Electrical Engineering, Or Related Field, 5+ Years Systems Programming Experience, and 3+ Years GPU Optimization Focus. Important skills include Nsight Systems, nvprof, CUDA Programming, GPU Architecture, Python, and C++.
Skills & qualifications
Skills
Qualifications
Full job description
We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.
The Role
You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.
Key Responsibilities
-
Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation
-
Write high-performance CUDA and Triton kernels for critical model operations
-
Optimize cold start latency from seconds to milliseconds for our serving infrastructure
-
Tune memory access patterns, kernel fusion, and GPU utilization
-
Collaborate with ML engineers to optimize model implementations
-
Debug performance issues across the full stack from application to hardware
-
Implement custom memory pooling and allocation strategies
-
Share optimization techniques and build performance culture across teams
Qualifications
-
Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
-
5+ years systems programming experience with 3+ years focused on GPU optimization
-
Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)
-
Strong CUDA programming skills with production kernel development
-
Deep understanding of GPU architecture (memory hierarchy, SMs, warps)
-
Track record of achieving significant performance improvements (5-10x)
-
Experience with Python and C++ in production environments
We Value
-
Experience with Triton kernel development
-
Knowledge of CUTLASS or similar high-performance libraries
-
Background in ML-specific optimizations (attention, transformers)
-
RDMA/InfiniBand optimization experience
-
Contributions to GPU libraries or frameworks
-
Low-level debugging skills (PTX/SASS reading)
Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish .
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Software Engineer - Planner GPU ComputeZoox · Boston, MA (Hybrid) · $237–298K/yrPosted 1w agoPosted 1w agoGPU Kernel EngineerSciforium · San Francisco, CA (Hybrid) · $190–250K/yrPosted 4 days agoPosted 4 days ago
Senior ML Accelerator Engineer - GPUGeneral Motors · Sunnyvale, CA (Hybrid) · $170–258K/yrPosted 4 days agoPosted 4 days agoSenior Performance EngineerCrusoe · San Francisco, CA · up to $205K/yrPosted 4 days agoPosted 4 days ago
Rendering Systems Engineer - Simulation & Synthetic DataWorld Labs · San Francisco, CA · $250–325K/yrPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.