Meta logo

Research Scientist, Multi-Modal Human Understanding

Meta

Pittsburgh, PAJob$122–181K/yrSeen 3w agoSeen in employer's feed 1 day ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$122–181K/yr
Location
Pittsburgh, PA
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Requirements

Credentials this posting asks for.

Bachelor's degree

Job overview

Meta is seeking a Research Scientist to advance multi‑modal AI technologies for human understanding and synthesis, developing vision‑language and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior and interaction at scale.

Skills & qualifications

RequiredNice to have

Skills

Multi-Modal ArchitecturesVision-Language ModelsVideo Foundation ModelsPyTorchTransformer-Based ArchitecturesPythonQuantitative AnalysisGenerative Models

Qualifications

Bachelor's Degree in Computer Science or Computer Engineering or Relevant Technical Field or Equivalent Practical Experience2+ Years of Multi-Modal AI Research Experience2+ Years of Experience Implementing and Training Large-Scale Neural Networks Using PyTorchExperience Designing and Executing Experiments to Evaluate Multi-Modal Model PerformanceExperience Writing Production-Quality or Research-Quality Code in Python for Multi-Modal AI ApplicationsExperience Developing or Fine-Tuning Vision-Language Models for Human Understanding TasksExperience With Video Foundation Models, Temporal Transformers, or Large-Scale Video PretrainingTrack Record of Contributing to Published Multi-Modal AI Research at Venues Such as CVPR, ICCV, or NeurIPSExperience With Generative Models for Human Synthesis Including Diffusion Models, GANs, or Autoregressive Models for Video or Motion Generation

Full job description

Summary:

Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.

Required Skills:

Research Scientist, Multi-Modal Human Understanding Responsibilities:

  1. Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding

  2. Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding

  3. Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis

  4. Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation

  5. Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization

  6. Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation

Minimum Qualifications:

Minimum Qualifications:

  1. Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta

  2. 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems

  3. 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures

  4. Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks

  5. Experience writing production-quality or research-quality code in Python for multi-modal AI applications

Preferred Qualifications:

Preferred Qualifications:

  1. Experience developing or fine-tuning Vision-Language Models for human understanding tasks

  2. Experience with video foundation models, temporal transformers, or large-scale video pretraining

  3. Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS

  4. Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation

Public Compensation:

$122,000/year to $181,000/year + bonus + equity + benefits

Industry: Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.