Member of Technical Staff - Infrastructure
San Francisco, CAFull-timePosted 3mo agoStill listed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Gimlet Labs seeks an Infrastructure Platform Engineer to build systems that transform heterogeneous accelerator hardware into reliable production infrastructure for its AI neocloud. The role involves deploying and operating clusters, automating hardware provisioning, enhancing scheduling and observability, and collaborating across distributed systems, runtime, compiler, networking, and hardware teams.
Skills & qualifications
Skills
Qualifications
Full job description
About us Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference. We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it. We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware. About the role As an Infrastructure Platform Engineer, you will build the systems that turn heterogeneous accelerator hardware into reliable production infrastructure for Gimlet's AI cloud. Gimlet's fleet spans hardware with different architectures, software stacks, operational characteristics, and failure modes. Your work will determine how new hardware is brought online, how clusters are provisioned and operated, and how production inference systems remain reliable as the fleet scales. You will work across bare metal, Linux, Kubernetes, cluster scheduling, observability, and automation. You will build systems that abstract differences between accelerator architectures, make new hardware production-ready, and improve the reliability and operability of Gimlet’s infrastructure. What success looks like In your first 12–18 months, you will:
-
Deploy and operate production clusters across different accelerator architectures
-
Automate hardware provisioning, validation, upgrades, and fleet lifecycle management
-
Improve cluster scheduling, resource utilization, isolation, and capacity management
-
Build observable infrastructure that enables faster debugging, incident response, and recovery
-
Partner across distributed systems, runtime, compiler, networking, and hardware teams to bring new accelerators into production
You may be a good fit if you have
-
Experience in infrastructure, cluster engineering, platform engineering, SRE, or HPC
-
Strong Linux systems knowledge and production debugging experience
-
Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
-
Experience automating infrastructure with Python, Go, Terraform, Ansible, or similar tools
-
Experience with GPU or accelerator infrastructure, including drivers, firmware, or CUDA/ROCm
-
The ability to build systems that are observable, recoverable, and reliable in production
-
A bachelor’s degree in a relevant field or equivalent practical experience
Strong candidates may also have
-
Experience building or operating AI inference, training, HPC, or neocloud infrastructure
-
Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up
-
Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting
-
Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks
-
Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling
-
Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators
Why join now? Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.
-
Solve hard problems.
-
Own meaningful work.
-
Build for production.
-
Help define what’s next.
Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Member of Technical Staff, Infrastructure EngineerVapi · San Francisco, CA (Hybrid) · $280–314K/yrPosted 3w agoPosted 3w ago
Member of Technical Staff (General Software Engineer, Infrastructure)Perplexity · San Francisco, CA (Hybrid) · $220–405K/yrPosted 4w agoPosted 4w agoMember of Technical Staff - ML Infrastructure Engineer, Post-trainingPreference Model · San Francisco, CA · $200–350K/yrPosted 3w agoPosted 3w ago
Member of Technical Staff, Infrastructure SecurityParallel · Bay Area, CA · $150–300K/yrPosted 2w agoPosted 2w agoMember of Technical Staff, Systems Infrastructure (2026 PhD New Grad)Fireworks · San Mateo, CA · $200–230K/yrPosted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.