Chai Discovery logo

Research Engineer - ML Infrastructure

Chai Discovery

San Francisco, CAFull-timePosted 6 days agoStill listed 5 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Chai Discovery builds a design suite for molecules, training frontier AI models that understand biochemical structure. The role focuses on making these models performant, resource‑efficient, and reliable at scale by developing core frameworks for training and evaluation in partnership with researchers and engineers.

Skills & qualifications

RequiredNice to have

Skills

PythonPyTorchJAXGPU ClustersParallelismQuantizationCustom KernelsFault ToleranceCheckpointingDeterministic OrchestrationProfiling

Full job description

About Chai Discovery Chai builds the design suite for molecules. We train frontier models that learn the underlying foundations of biochemical structure and interaction, so scientists can move faster and pursue targets that other methods cannot reach. AI is reinventing life sciences the same way it reinvented software engineering, and Chai is at the forefront of this shift. Leading pharmaceutical companies like Eli Lilly, Pfizer, and Novartis are adopting our platform to power their drug discovery programs. We value diverse perspectives and are ready to find greatness in unexpected places. About the role Make our models performant, resource efficient and reliable at scale by developing the core frameworks for model training and evaluation, in close partnerships with fellow researchers and engineers.

  • Build our training stack across model, layer, and kernel levels; optimize workloads through parallelism, quantization, and custom kernels.

  • Profile end-to-end training runs on large GPU clusters; eliminate bottlenecks and failures; monitor throughput, utilization, and uptime.

  • Ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs.

  • Make our ML training stack maximally reliable: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.

Chai's models are moving beyond protein structure prediction into real-world therapeutic engineering. This is a chance to push the frontier of AI drug design, working alongside a rigorous and craft-obsessed team. About you Ideal backgrounds include deep industry experience working with top AI/ML teams on the kinds of problems and systems we describe above—with strong software system design skills, proficiency in Python, and Pytorch or JAX fluency. We look for technical spikes where you have gone deep and demonstrated exceptional impact on real-world problems and systems. We offer The opportunity to work at the vanguard of AI research and frontier biology, with world-class people, on a mission that matters. We protect & promote a culture of high velocity and ownership. We compensate our team accordingly.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.