Research Scientist, Multi-Modal Understanding & Synthesis
Pittsburgh, PAJob$184–257K/yrSeen 1w agoSeen in employer's feed 1 day ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Pittsburgh, PA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
Meta is seeking a Research Scientist to drive foundational research in multi-modal understanding, synthesis, and world models, advancing AI systems that perceive, reason across, and generate content across vision, language, audio, and other modalities.
Skills & qualifications
Skills
Qualifications
Full job description
Summary:
Meta is seeking a Research Scientist to drive foundational research in multi-modal understanding, synthesis, and world models. In this role, you will advance the state of the art in building AI systems that perceive, reason across, and generate content spanning vision, language, audio, and other modalities. You will develop world models that learn rich internal representations of human behavior, enabling prediction, planning, and simulation. Collaborating with world-class researchers and engineers, you will define research directions, publish influential work, and translate breakthroughs into technologies that power Meta's next-generation AI products.
Required Skills:
Research Scientist, Multi-Modal Understanding & Synthesis Responsibilities:
-
Lead original research in multi-modal learning, developing architectures and algorithms that unify understanding and generation across vision, language, audio, and other modalities
-
Design and build world models that learn predictive representations of human behavior, supporting capabilities such as simulation, planning, and reasoning
-
Drive end-to-end research projects from problem formulation and dataset curation through model development, evaluation, and integration into real-time prototypes
-
Develop novel approaches for multi-modal synthesis, enabling coherent generation of images, video, text, and audio from unified representations
-
Establish rigorous evaluation frameworks, benchmarks, and metrics to measure progress in multi-modal reasoning and world modeling
-
Mentor other engineers and researchers on the team, providing technical guidance on multi-modal architectures, generative models, and research best practices
-
Publish research findings at top-tier peer-reviewed venues such as NeurIPS, ICLR, and CVPR to advance the broader scientific community
Minimum Qualifications:
Minimum Qualifications:
-
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
-
PhD in Machine Learning, Computer Vision, Natural Language Processing, or a closely related field
-
6+ years of experience conducting AI research in multi-modal learning, generative models, or world models, including experience leading major research initiatives from conception through publication or production deployment
-
Experience implementing and evaluating multi-modal systems using deep learning frameworks such as PyTorch or TensorFlow, with proficiency in Python
-
Experience publishing original research in peer-reviewed machine learning or AI venues
-
Experience driving cross-functional technical decisions and communicating research findings and trade-offs to both research and engineering audiences through written documents and presentations
-
Experience with techniques spanning multiple modalities such as vision-language models, multi-modal transformers, or cross-modal representation learning
Preferred Qualifications:
Preferred Qualifications:
-
Experience developing large-scale multi-modal foundation models or vision-language models
-
Experience with world models, predictive learning, or model-based reinforcement learning for planning and reasoning
-
First-author publications at top-tier venues such as NeurIPS, ICLR, or CVPR demonstrating contributions to multi-modal learning, generative models, or world models
Public Compensation:
$184,000/year to $257,000/year + bonus + equity + benefits
Industry: Internet
Equal Opportunity:
Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior AI Engineer, Security Infrastructure Air Company · Pittsburgh, PAPosted 3w agoPosted 3w ago
Interdisciplinary General Engineer (Recent Graduate)Department of Energy Headquarters · Albany, OR · $56–86K/yrPosted todayPosted today
Interdisciplinary Physical Science (Recent Graduate)Department of Energy Headquarters · Albany, OR · $50–86K/yrPosted todayPosted today
Interdisciplinary Physical Scientist/ChemistMine Safety and Health Administration · South Park, PA · $52–100K/yrPosted 1 day agoPosted 1 day agoResearch Scientist, Multi-Modal Human UnderstandingMeta · Pittsburgh, PA · $122–181K/yr
You've read the whole posting — now see how you match it.