Causal Labs logo

Member of Technical Staff - ML Infra

Causal Labs

San Francisco, CAFull-timePosted 1y agoStill listed 2 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Causal Labs is hiring a Member of Technical Staff - ML Infra. Causal Labs is seeking a Member of Technical Staff for ML Infrastructure. This role involves designing, deploying, and maintaining large distributed ML training and inference clusters. The ideal candidate will develop efficient, scalable end-to-end pipelines for petabyte-scale datasets and model training, research various training approaches, and optimize GPU operations. They should possess a relentless approach to problem-solving and a rapid execution ability.

Key focus areas include Design, deploy, and maintain large distributed ML training and inference clusters, Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training, and Research and test various training approaches including parallelization techniques.

Preferred (not required): Distributed ML Training, Distributed ML Inference, ML Lifecycle Management, and Parallelization Techniques.

Skills & qualifications

RequiredNice to have

Skills

Distributed ML TrainingDistributed ML InferenceML Lifecycle ManagementParallelization TechniquesNumerical Precision Trade-OffsGPU Operations OptimizationProblem-SolvingRapid ExecutionQuick LearningOptimizing Training WorkloadsOptimizing Inference WorkloadsFSDPDeepSpeedGCPAWSAzureKubernetesDockerDistributed Task Management SystemsScalable Model Serving ArchitecturesScalable Model Deployment ArchitecturesMonitoring Best Practices for ML Systems

Full job description

Responsibilities

  • Design, deploy, and maintain large distributed ML training and inference clusters

  • Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle

  • Research and test various training approaches including parallelization techniques and numerical precision trade-offs across different model scales

  • Analyze, profile and debug low-level GPU operations to optimize performance

  • Stay up-to-date on research to bring new ideas to work

What we’re looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

  • Strong grasp of state-of-the-art techniques for optimizing training and inference workloads

  • Demonstrated proficiency with distributed training frameworks (e.g. FSDP, DeepSpeed) to train large foundation models

  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings

  • Familiarity with containerization and orchestration frameworks (e.g., Kubernetes, Docker)

  • Background working on distributed task management systems and scalable model serving & deployment architectures

  • Understanding of monitoring, logging, observability, and version control best practices for ML systems

You don’t have to meet every single requirement above.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.