Between logo

Machine Learning Engineer, Speech/Audio

Between

San Francisco, CA · HybridFull-timePosted 2mo agoStill listed 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
San Francisco, CAHybrid
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Between is hiring a Machine Learning Engineer, Speech/Audio. Between is seeking a Machine Learning Engineer specializing in Speech/Audio to own the audio modeling stack from raw signal to production inference. This role involves shipping models that determine, in real-time, who a voice agent is being spoken to. The engineer will work with founders and early customers to shape model development.

Key focus areas include Own the audio model lifecycle end to end, Build and improve models for addressee detection, and Turn research prototypes into fast, dependable inference.

Important skills include Audio Modeling Stack Ownership, Model Lifecycle Management, Addressee Detection Model Building, Research Prototype Conversion To Inference, Evaluation Harness Design, and Dataset Design. Preferred (not required): Real Time Audio, Source Separation, and On-Device Inference.

Skills & qualifications

RequiredNice to have

Skills

Audio Modeling Stack OwnershipModel Lifecycle ManagementAddressee Detection Model BuildingResearch Prototype Conversion to InferenceEvaluation Harness DesignDataset DesignApplied Machine LearningAudio ModelsSpeech ModelsSignal-Heavy ModelsModel Ownership From Data Collection to ProductionPyTorchMeasurement DisciplineBias Toward ShippingSmallest Experiment DesignSystems SenseReal Time AudioSource SeparationOn-Device Inference

Full job description

Careers

Machine Learning Engineer, Speech/Audio San Francisco · Full-time · Machine Learning

You will own the audio modeling stack end to end, from raw signal to production inference. You are the person who ships the models that decide, in real time, who a voice agent is being spoken to.

Apply What you will do

  • Own the audio model lifecycle end to end: data, features, architecture, training, evaluation, and the inference path that runs in production.

  • Build and improve the models behind addressee detection, deciding whether speech is meant for the agent or for someone else in the room.

  • Turn research prototypes into fast, dependable inference that holds up under real world noise, overlap, and accents.

  • Design the evaluation harness and datasets that tell us, honestly, whether a change made the product better.

  • Work directly with founders and early customers to shape what the model needs to do next. What we are looking for

  • Strong applied machine learning experience with audio, speech, or other signal-heavy models.

  • Comfort owning a model from data collection through to something running in production, not just a notebook.

  • Fluency with modern deep learning tooling such as PyTorch, and the discipline to measure before and after every change.

  • A bias toward shipping, and toward the smallest experiment that answers the question.

  • Enough systems sense to care about latency, memory, and how a model behaves under load.

  • Bonus: experience with real time audio, source separation, or on-device inference. About attention labs

attention labs is early: a small team defining a new category at the intersection of speech, cognitive neuroscience, and machine learning. We work from San Francisco, Toronto, and Memphis, and we are remote-friendly for the right person.

Sound like you?

Send a note and tell us why this role fits. We read every application.

Apply for this role

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.