Bland AI logo

Machine Learning Researcher, Multimodal LLMs

Bland AI

San Francisco, CAHybridFull-time$180–260K/yrPosted 4mo agoVerified open 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$180–260K/yr
Location
San Francisco, CAHybrid
Schedule
Full-time
Work Authorization
Not specified

Job overview

Bland AI is hiring a Machine Learning Researcher, Multimodal LLMs. Bland AI seeks a Machine Learning Researcher to develop next‑generation multimodal LLM stack, integrating speech, text, tools, and real‑time reasoning, building conversational AI models from idea through production for millions of calls daily. The role involves defining how agents listen, think, and act in real time, and collaborating across product and engineering teams.

Key focus areas include Develop multimodal LLM stack integrating speech, text, tools, and real‑time reasoning, Build conversational AI models from research to production serving millions of calls, and Define agent listening, thinking, and acting mechanisms in real time.

Important skills include LLMs, Multimodal Models, Speech-Language Systems, Prompting, Fine-Tuning, and Alignment Techniques. Preferred (not required): Real-Time Voice Systems, Conversational AI, Tool-Using Agents, and Agent Frameworks.

Skills & qualifications

RequiredNice to have

Skills

LLMsMultimodal ModelsSpeech-Language SystemsPromptingFine-TuningAlignment TechniquesNeural Audio CodecsModern Multimodal LLM TechniquesFast Experimental LoopDesign ExperimentsProduct IntuitionTranslate Modeling IdeasBuilder MentalityOwnership From Research Through DeploymentThrive in Ambiguous EnvironmentsThrive in Fast-Moving EnvironmentsCare About ImpactThink in SystemsObsess Over LatencyObsess Over CorrectnessObsess Over Real-World BehaviorComfortable Discarding IdeasPush Toward Simple AbstractionsReal-Time Voice SystemsConversational AITool-Using AgentsAgent FrameworksMultimodal DatasetsLLM Research ContributionsSpeech-Related Research ContributionsOpen Source Contributions

Benefits

Medical Insurance
Dental Insurance
Vision Insurance

Full job description

Machine Learning Researcher, Multimodal LLMs Location: San Francisco, CA or Remote

About Bland At Bland.com , our mission is to empower enterprises to build AI phone agents at scale. Voice is quickly becoming the primary interface between businesses and their customers, and we are building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human.

We’ve raised $100M from leading investors including Emergence Capital, Scale Venture Partners, Y Combinator, and founders of Twilio, Affirm, and ElevenLabs.

The Role We are looking for someone to contribute to the development of our next-generation multimodal LLM stack, combining speech, text, tools, and real-time reasoning into a single unified system. You’ll be responsible for building industry-leading conversational AI models that power Bland's agent, and taking them all the way from idea to production.

At Bland, we're not just thinking about text modeling. You will define how our agents listen, think, and act in real time , integrating streaming audio, tool execution, and dynamic context into a single coherent system. You will take ideas from research through production systems serving millions of calls per day.

What Makes You a Great Fit

Strong LLM / Multimodal Background

  • Experience with LLMs, multimodal models, or speech-language systems

  • Deep understanding of prompting, fine-tuning, and alignment techniques

  • Familiarity with neural audio codecs and modern multimodal LLM techniques

Fast Experimental Loop

  • You can go from idea → dataset → experiment → conclusion in days

  • You know how to design experiments that actually answer the question

Product Intuition

  • Strong sense for what makes an interaction feel natural vs robotic

  • Ability to translate abstract modeling ideas into user-facing improvements

Builder Mentality

  • You take ownership from research through deployment

  • You thrive in ambiguous, fast-moving environments

  • You care about impact, not just elegance

How You Show Up

  • You think in systems, not just models

  • You obsess over latency, correctness, and real-world behavior

  • You are comfortable discarding ideas quickly when data disagrees

  • You push toward simple abstractions for complex problems

Bonus Points

  • Experience with real-time voice systems or conversational AI

  • Background in tool-using agents or agent frameworks

  • Experience with multimodal datasets (audio + text + actions)

  • Contributions to LLM or speech-related research or open source

Compensation & Benefits

  • Competitive salary: $180,000 – $260,000

  • Meaningful equity

  • Full healthcare, dental, vision

  • Office in Levi's Plaza, SF

  • High autonomy, high impact

You've read the whole posting — now see how you match it.