Machine Learning Research Intern, Audio
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
The role is a Machine Learning Research Intern focused on audio, where the intern will own a research project across the voice stack, work with the research team on real‑world problems, and aim for results that can be shipped or published.
Skills & qualifications
Skills
Qualifications
Full job description
The Role: Machine Learning Research Intern, Audio As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy. We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.
What You Will Do Own a research question end to end
-
Take one well-scoped problem from literature review through implementation, experimentation, and results.
-
Design ablations that isolate what actually caused an improvement.
-
Present your findings to the research team and defend the methodology.
Work on real systems
-
Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
-
Use our distributed GPU infrastructure rather than toy-scale setups.
-
Where the result warrants it, work with engineers to move it toward production.
Choose your depth Depending on your background and interests, your project may focus on:
-
Expressive and controllable text-to-speech, including prosody and emotion modeling
-
Neural audio codecs and discrete or continuous speech representations
-
ASR robustness for telephony, accents, and code switching
-
Real-time and streaming inference under latency constraints
-
Full-duplex conversation and turn-taking dynamics
What Makes You a Great Fit Research foundations
-
Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.
-
Comfortable reading a paper and reimplementing it without hand-holding.
-
Experience with self-supervised, generative, or multimodal modeling.
Audio or speech grounding
-
Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.
-
Strong intuition for audio quality and what makes synthetic speech sound wrong.
-
Prior publications or open source contributions in speech or language AI are a strong signal, though not required.
Engineering ability
-
Fluent in PyTorch and comfortable in a real codebase.
-
Able to run your own experiments on GPU clusters without waiting to be unblocked.
How You Show Up
-
You identify the single experiment that validates an idea in days, not months.
-
You measure everything and let data drive decisions.
-
You are honest about negative results, because they are how we narrow the search.
-
You are obsessed with making voice agents sound truly human.
-
You use AI tools aggressively to amplify your own impact.
Benefits
-
Competitive intern compensation
-
Mentorship from researchers working on frontier voice AI
-
Every tool you need to succeed
-
Beautiful office in Levi's Plaza, SF with rooftop views
-
A real shot at a return offer
You've read the whole posting — now see how you match it.