Ultravox.ai logo

Research Engineer / Scientist, Foundational Multimodal Models

Ultravox.ai

Seattle, WAJobPosted 1y agoStill listed 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Seattle, WA
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Ultravox.ai is a Seattle‑based AI startup building open‑source speech‑to‑speech models. The company has raised seed funding and focuses on real‑time voice AI, seeking a Research Engineer & Scientist to lead development of foundational multimodal models and collaborate across research and product teams.

Skills & qualifications

RequiredNice to have

Skills

Architectural DesignProject CollaborationData AnalysisAI ResearchLarge Language ModelsSpeech ModelsMultimodal ModelsPythonPyTorchPre‑TrainingPost‑TrainingReal‑Time Voice AIData Flywheel DevelopmentModel Quality MeasurementModel OptimizationModel Deployment

Qualifications

AI Researcher Track RecordExperience With Large Language ModelsExperience With Speech ModelsExperience With Multimodal ModelsStrong Python ExperienceAbility to Roll Up SleevesGreat Communicator and Team Player

Benefits

Paid Time Off
Medical Insurance
401(k) Match

Full job description

About Fixie We’re a Seattle-based AI startup (with support for working remotely). We’ve raised $17M in seed funding. Our vision is simple: build artificial intelligences that can communicate as naturally as humans. We’re a small team of researchers and engineers with a deep focus in speech and real-time technologies. Our core model, Ultravox, is open-source. We also build a serving stack that’s optimized for very low-latency interactions. The Role As a Research Engineer & Scientist, you will lead the effort to develop next-generation foundational multimodal models that power Ultravox, our open-source speech-to-speech model. What you’ll do

  • Lead critical research on architectural design, pre-training, and post-training of foundational multimodal models to develop real-time voice AI.

  • Collaborate with a team of researchers and engineers to develop foundational multimodal models with comprehensive capabilities in speech understanding, speech generation, and full-duplex real-time communication.

  • Develop novel models based on public and proprietary data sources.

  • Build tools to improve our data flywheel and measure model quality.

  • Drive the optimization and deployment of AI models for real-world applications in partnership with engineering and product teams. Things we’re looking for

  • An incredibly strong AI researcher with a track record of contributions to AI research, systems, and products.

  • Experience with large language models, speech models, and multimodal models.

  • Strong experience in Python and, ideally, PyTorch.

  • Ability to roll up your sleeves and get things done.

  • A great communicator and team player. Benefits

  • Generous equity package

  • Unlimited PTO (take time when you need it)

  • Top-of-market salary

  • Great healthcare

  • 401k with match Apply

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.