Senior Site Reliability Engineer
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
Luma AI is hiring a Senior Site Reliability Engineer. Luma AI is seeking a Senior Site Reliability Engineer to own and maintain their GPU infrastructure, which includes thousands of NVIDIA and AMD GPUs across on-premise and multi-cloud environments. This hands-on role involves ensuring the reliability and speed of training and inference clusters, redesigning systems for scalability, and debugging complex, low-level issues. The ideal candidate thrives in a fast-paced, less-structured environment and possesses deep Linux expertise.
Key focus areas include Take end-to-end ownership of production GPU clusters, Join critical re-architecture sessions to redesign systems, and Tune Linux performance deeply, at the OS and kernel level.
Successful candidates bring 5+ Years SRE Experience, 5+ Years Production Engineer Experience, and 5+ Years Infrastructure Engineer Experience. Important skills include AWS, OCI, Linux, GPU Infrastructure, Networking, and Kernel-Level Debugging. Preferred (not required): NVIDIA, AMD, DCGM, and ROCm.
Skills & qualifications
Skills
Qualifications
Full job description
Team: Infra Reliability · SF Bay Area / Remote (US)
You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference clusters reliable and fast, and you help redesign them for the next level of scale.
This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it.
What You'll Own
-
Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
-
Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
-
Tune Linux performance deeply, at the OS and kernel level.
-
Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
-
Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
-
Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.
First 90 Days
One way the first 90 could unfold.
-
Days 1–30 — Immerse & Diagnose: Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
-
Days 30–60 — Ship & Validate: Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
-
Days 60–90 — Scale & Systemize: Contribute to the next-gen re-architecture and harden security and compliance practices.
What You Bring
-
5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
-
Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
-
Working experience with Terraform, Airflow, and Ray.
-
Strong experience with AWS or OCI.
-
Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
-
Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
-
Comfort in a less-structured, fast-paced environment.
Nice to Have
-
Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
-
Experience managing large-scale GPU clusters for AI/ML training or inference.
-
Familiarity with Kubernetes or orchestration frameworks like Ray.
-
Deep expertise in data pipelines and infrastructure.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
You've read the whole posting — now see how you match it.