Pocket logo

Research Engineer - Voice Intelligence

Pocket

San Francisco, CAFull-timeSeen 1mo agoStill listed 2 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Pocket is building a category-defining platform for ambient intelligence. Voice is the highest-signal interface they have, and they are seeking a research‑minded engineer to push what’s possible with voice end‑to‑end, from capture and understanding to real‑time experiences and evaluation.

Skills & qualifications

RequiredNice to have

Skills

DiarizationVoice Activity DetectionNoise RobustnessTranscription QualitySpeaker IdentificationPost‑ProcessingStreaming Audio PipelinesReal‑Time SystemsLLM Post‑ProcessingDataset CurationLabeling ToolingQuality Assurance

Qualifications

6–10+ Years Production Systems Experience

Full job description

Pocket is building a category-defining platform for ambient intelligence. Voice is the highest-signal interface we have — and we’re looking for a research-minded engineer to push what’s possible with voice end-to-end: capture → understanding → real-time experiences → evaluation. What you'll own

  • Invent and ship improvements to our voice stack: diarization, VAD, noise robustness, transcription quality, speaker ID, and post-processing.
  • Build new voice-powered product capabilities: better memory, better meeting/voice summaries, better action extraction, better personalization.
  • Run tight research loops: define metrics, build eval sets, iterate on models/algorithms, and productionize what works.
  • Improve real-time performance and reliability: streaming pipelines, latency budgets, fallbacks, and graceful degradation.
  • Partner closely with product + design to translate voice capabilities into features people feel immediately.

What success looks like

  • Step-function improvements in voice quality and perceived intelligence (measured and felt).
  • A clear evaluation harness (offline + online) that prevents regressions and accelerates iteration.
  • Voice features that ship reliably at scale with predictable latency.

What we're looking for

  • Have 6–10+ years of experience shipping production systems with strong technical ownership.
  • Are strong in applied ML/research engineering and can turn prototypes into robust product.
  • Understand audio/voice fundamentals (signal processing basics helpful) and modern model/eval workflows.
  • Care deeply about performance and reliability.

Nice-to-haves

  • Experience with streaming audio pipelines and real-time systems.
  • Experience with LLM post-processing, structured extraction, and evaluation.
  • Experience building tooling for labeling, dataset curation, and QA.

What we offer

  • Work directly with us and learn fast
  • Direct impact on how the company operates day to day
  • High-trust, high-responsibility environment
  • Competitive compensation

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.