Tavus logo

Multimodal AI Model Optimization Research Engineer

Tavus

Remote · USJobPosted 6mo agoStill listed 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Remote · US
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Tavus is building the human layer of AI to enable natural human‑AI interaction across industries such as healthcare, recruiting, sales, and education, focusing on multimodal models that see, hear, and communicate. The role joins the core AI team to accelerate research models into fast, efficient, production‑ready systems.

Skills & qualifications

RequiredNice to have

Skills

BenchmarkingGPULanguage ModelsLow LatencyPerformance TuningCUDAPyTorchMachine LearningArtificial IntelligenceResearchWebRTCDecision TreesStreamingModel OptimizationModel CompressionKnowledge DistillationPruningQuantizationMixed PrecisionLow‑Rank AdaptersInference PerformanceGPU FundamentalsPythonLarge ModelsCloud EnvironmentsML PapersCommunicationDiffusion ModelsVideo/Audio Generative ModelsLarge Language ModelsReal‑Time Streaming SystemsStreaming TTS/VideoTensorRTONNX RuntimeTVMTritonXLACustom Triton/CUDA KernelsLow‑Level Performance TuningExperiment TrackingProfilingResearch EngineeringApplied Science

Qualifications

Strong Experience in Deep Learning Using PyTorchHands‑on Experience With Model Optimization and CompressionUnderstanding of Efficient Architectures Such as Low‑Rank AdaptersStrong Understanding of Inference Performance and GPU FundamentalsStrong Python Coding SkillsExperience Working With Large Models and Datasets in Cloud EnvironmentsAbility to Read ML Papers and Reproduce ResultsClear Communication and Collaboration SkillsOptimization of Diffusion Models or Video/Audio Generative Models or Large Language ModelsExperience With Real‑Time or Streaming Systems Such as WebRTC or Streaming TTS/VideoFamiliarity With TensorRT or ONNX Runtime or TVM or Triton or XLAExperience Writing Custom Triton/CUDA Kernels or Low‑Level Performance Tuning

Benefits

Paid Time Off
Medical Insurance

Full job description

About Us

At Tavus, we're building the human layer of AI. Our mission is to make human-AI interaction as natural as face-to-face interaction, enabling the human touch where it has been previously unscalable. We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models power everything from text-to-video AI avatars to real-time conversational video experiences across industries like healthcare, recruiting, sales, and education.

By enabling AI to see, hear, and communicate with human-like authenticity, we're creating the foundation for the next generation of AI employees, assistants, and companions.

We are a Series B company backed by top investors, including Sequoia, Y Combinator, and Scale VC. Join us in driving the future of human-AI interaction.

The Role We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team. Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks. We’re moving fast and looking for people who can help pave the path.

Your Mission

  • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization

  • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality

  • Partner closely with researchers and engineers to turn new ideas into deployable systems

Requirements

  • Strong experience in deep learning using PyTorch

  • Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision

  • Understanding of efficient architectures such as low-rank adapters

  • Strong understanding of inference performance and GPU/accelerator fundamentals

  • Strong Python coding skills and reliable research engineering practices

  • Experience working with large models and datasets in cloud environments

  • Ability to read ML papers, reproduce results, and adapt ideas

  • Clear communication and collaboration skills

Preferred Experience

  • Optimization of diffusion models, video/audio generative models, or large language models

  • Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)

  • Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA

  • Experience writing custom Triton/CUDA kernels or low-level performance tuning

  • Experience with experiment tracking, benchmarking, and profiling at scale

  • Prior experience in research engineering or applied science roles

Location This position is preferably hybrid in San Francisco, with relocation support offered. Remote candidates are also considered.

Benefits When you join Tavus, you’re joining a family. We offer flexible work schedules, unlimited PTO, competitive healthcare and gear stipends, and a collaborative environment focused on learning and impact.

Culture & Diversity We are not looking for cultural fits — we are looking for culture creators. Diversity drives our success, and we combine varied backgrounds, skills, and perspectives to build the best experiences for our clients..

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.