Liquid AI logo

Member of Technical Staff - Multi-Modal, Vision

Liquid AI

San Francisco, CAHybridFull-timeNo compensation foundPosted 10mo agoChecked 1w ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
San Francisco, CAHybrid
Schedule
Full-time
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Master's degree

Job overview

Liquid AI is hiring a Member of Technical Staff - Multi-Modal, Vision. Liquid AI, spun out of MIT CSAIL, develops general-purpose AI systems for efficient deployment across various targets, from data centers to on-device hardware. The company is rapidly expanding and seeks exceptional individuals. The VLM team focuses on building vision-language models that operate on-device with strict latency and memory constraints, while maintaining high quality. This role involves owning the full VLM pipeline, from research to deployment.

Key focus areas include Lead a new model capability end-to-end from task specification through data curation, Improve visual reasoning through reinforcement learning and preference optimization methods, and Push the quality-efficiency frontier on token efficiency via encoder/connector design.

Successful candidates bring M.S. Or Ph.D. In Computer Science, Mathematics, Or Related Field and Equivalent Industry Experience. Important skills include Training VLMs, Evaluating VLMs, Experimental Rigor, Scalable Implementations, Refine Hypotheses, and Iterate Hypotheses. Preferred (not required): Multimodal Training Pipelines, Optimizing Multimodal Data Pipelines, Distributed Training, and DeepSpeed.

Skills & qualifications

RequiredNice to have

Skills

Training VLMsEvaluating VLMsExperimental RigorScalable ImplementationsRefine HypothesesIterate HypothesesPythonDeep Learning FrameworkMultimodal Training PipelinesOptimizing Multimodal Data PipelinesDistributed TrainingDeepSpeedFSDPMegatron-LMMultimodal Post-TrainingSFTPreference OptimizationRL-Style MethodsDataset DesignData Quality ExpertiseQuality AssessmentDiversity AssessmentLong-Tail MiningOpen-Source ContributionsGitHubHugging FacePublished ResearchNeurIPSICMLCVPRECCVICLRACLComputer VisionVisual Representation Learning

Qualifications

M.S. Or Ph.D. In Computer Science, Mathematics, or Related FieldEquivalent Industry Experience

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match
Paid Time Off

Full job description

ABOUT LIQUID AI

Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.

THE OPPORTUNITY

The VLM team builds vision-language models that run on-device, under tight latency and memory constraints, without sacrificing quality. We have released four best-in-class models and we're just getting started.

This team owns the full VLM pipeline end-to-end: from researching new architectures and training algorithms through data curation, evaluation, and deployment. You'll join a focused, hands-on group that works directly on models and collaborates closely with our pretraining, post-training, and infrastructure teams. Success here is measured by the capability of the models we ship.

MINIMAL QUALIFICATIONS:

  • Hands-on experience in training or evaluating VLMs with demonstrated experimental rigor.

  • Ability to turn research ideas into scalable implementations, refine and iterate through hypotheses.

  • Proficiency in Python and at least one deep learning framework.

  • M.S. or Ph.D. in Computer Science, Mathematics, or a related field; or equivalent industry experience.

THIS ROLE IS FOR YOU IF YOU HAVE EXPERIENCE IN SOME OF THE FOLLOWING:

  • Building or optimizing multimodal training or data pipelines.

  • Experience with distributed training (DeepSpeed, FSDP, Megatron-LM, etc.).

  • Multimodal post-training experience (SFT, preference optimization, RL-style methods).

  • Dataset design and data quality expertise (quality and diversity assessment, long-tail mining).

  • Prior open-source contributions (code, data, models) on GitHub or Hugging Face.

  • Published research at top AI conferences (NeurIPS, ICML, CVPR, ECCV, ICLR, ACL, etc.).

  • Experience with computer vision or visual representation learning.

WHAT WORKING HERE MIGHT LOOK LIKE:

  • Lead a new model capability end-to-end from task spec through data curation, training recipe, ablations, evaluation, and into the final shipped model.

  • Improve visual reasoning through reinforcement learning and preference optimization methods.

  • Push the quality-efficiency frontier on token efficiency via encoder/connector design. Exemplary outcome: a connector that cuts vision tokens without quality loss.

WHAT SUCCESS LOOKS LIKE (YEAR ONE):

  • The VLM models we ship are state-of-the-art.

  • You own a major work-stream (for instance, video understanding, preference data quality, or encoder architecture) end-to-end.

  • At least one model has shipped to production with your direct contribution.

WHAT WE OFFER:

  • Full ownership: You own your work from architecture to deployment.

  • Compensation: Competitive base salary with equity in a unicorn-stage company

  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents

  • Financial: 401(k) matching up to 4% of base pay

  • Time Off: Unlimited PTO plus company-wide Refill Days throughout the year

You've read the whole posting — now see how you match it.