
HPC/ML Infrastructure Engineer
San Francisco, CAFull-timePosted 3mo agoStill listed 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near San Francisco, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
The team seeks an experienced HPC infrastructure engineer to lead the bring‑up, administration and operation of one of the world’s largest anime AI training clusters, acting as the bridge between researchers and GPU hardware while ensuring jobs run smoothly and systems stay reliable.
Skills & qualifications
Skills
Qualifications
Full job description
We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training. You may be a good fit if: You love anime and the anime aesthetic. This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems. You’re familiar with the modern HPC software landscape Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!) As well as the traditional sysadmin landscape Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools. You're not afraid of physical computers We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help. And you're comfortable working on small, fast-paced teams. We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world. We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Research Engineer - ML InfrastructureChai Discovery · San Francisco, CAPosted 1w agoPosted 1w ago
2027 Internship Onboard Infrastructure Engineer, ML InferenceBedrock Robotics · San Francisco, CAPosted 3w agoPosted 3w ago
ML EngineerMonarch · Emeryville, CAPosted 3w agoPosted 3w ago
Senior Software Engineer, ML Infra - Asset SafetyRoblox · San Mateo, CA · $279–329K/yrPosted 1w agoPosted 1w agoMember of Technical Staff - ML Infrastructure Engineer, Post-trainingPreference Model · San Francisco, CA · $200–350K/yrPosted 2w agoPosted 2w ago
You've read the whole posting — now see how you match it.