Between logo

Embedded Systems Engineer, On-Device Inference

Between

Remote · USFull-timePosted 2mo agoStill listed 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Remote · US
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Between is hiring an Embedded Systems Engineer, On-Device Inference. The Embedded Systems Engineer will bring addressee-detection models on-device, building efficient inference for regulated and OEM hardware where a GPU is not an option. This role involves optimizing inference with tools like ONNX, quantization, and careful profiling, working close to the metal on audio pipelines, memory, and latency budgets. The engineer will partner with research and OEM customers to fit models to real hardware and constraints, and build tooling and tests to ensure on-device behavior.

Key focus areas include Take models from the server to the device, running efficiently on constrained CPUs, Optimize inference with tools like ONNX, quantization, and careful profiling, and Work close to the metal on audio pipelines, memory, and latency budgets.

Important skills include C, C++, Python, Inference Runtimes, Correctness, and Reliability. Preferred (not required): ONNX, Quantization, Profiling, and Audio DSP.

Skills & qualifications

RequiredNice to have

Skills

CC++PythonInference RuntimesCorrectnessReliabilityOwning Problems End to EndONNXQuantizationProfilingAudio DSPReal Time ConstraintsShipping Software Into Physical Devices

Qualifications

Strong Systems ExperienceEmbedded ExperienceTrack Record of Making Models or Signal Code Run Fast on Constrained Hardware

Full job description

Careers

Embedded Systems Engineer, On-Device Inference Memphis · Full-time · Engineering

You will bring the addressee-detection model on-device, building efficient inference for regulated and OEM hardware where a GPU is not an option. This is the work that lets our models run where the audio actually happens.

Apply What you will do

  • Take our models from the server to the device, running efficiently on constrained CPUs.

  • Optimize inference with tools like ONNX, quantization, and careful profiling, without a GPU to lean on.

  • Work close to the metal on audio pipelines, memory, and latency budgets.

  • Partner with research and OEM customers to fit the model to real hardware and real constraints.

  • Build the tooling and tests that keep on-device behavior honest across devices. What we are looking for

  • Strong systems and embedded experience: C or C++, plus comfort in Python for tooling.

  • A track record of making models or signal code run fast on constrained hardware.

  • Familiarity with inference runtimes, quantization, and profiling for CPU targets.

  • Care for correctness and reliability in environments you cannot easily update.

  • Comfort owning a hard, specific problem end to end.

  • Bonus: experience with audio DSP, real time constraints, or shipping software into physical devices. About attention labs

attention labs is early: a small team defining a new category at the intersection of speech, cognitive neuroscience, and machine learning. We work from San Francisco, Toronto, and Memphis, and we are remote-friendly for the right person.

Sound like you?

Send a note and tell us why this role fits. We read every application.

Apply for this role

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.