Luma AI logo

Qualitative Evaluation Engineer

Luma AI

RemoteRemoteJobNo compensation foundPosted 3w agoVerified open 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
RemoteRemote
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Master's degree

Job overview

Luma AI is hiring a Qualitative Evaluation Engineer. The Qualitative Evaluation Engineer will own how Luma judges model quality beyond numeric metrics, building frameworks to assess believability, identity retention, and scene coherence, and translating insights into product improvements.

Key focus areas include Evaluate generative model performance across diverse tasks, prompts, and modalities, surfacing failure modes and edge cases, Build and maintain scalable, reusable qualitative evaluation frameworks, and Translate high-level product goals into concrete evaluative criteria.

Important skills include Cognitive Science, Design Research, Psychology, Communication Studies, Systems Thinking, and Written Communication. Preferred (not required): Motion, Visual Effects, Storytelling Pipelines, and Evaluating AI-Generated Media.

Skills & qualifications

RequiredNice to have

Skills

Cognitive ScienceDesign ResearchPsychologyCommunication StudiesSystems ThinkingWritten CommunicationCollaborationData CollectionEvaluate Generative Model PerformanceBuild Qualitative Evaluation FrameworksTranslate Product Goals Into Evaluative CriteriaLead Qualitative StudiesLead Side-by-Side ComparisonsLead Human-in-the-Loop EvaluationsTurn Nuanced Judgments Into Clear FeedbackWork With Technical Artists and EngineersCreative Workflows for Generative ModelsDefine Abstract Qualities in Evaluative TermsExcellent Written CommunicationSynthesize Nuanced Judgment Into Actionable InsightWorking Across Engineers, Researchers, and CreativesMotionVisual EffectsStorytelling PipelinesEvaluating AI-Generated MediaBuilding Internal Tools for Qualitative Data CollectionBuilding Internal Tools for ScoringPrompt EngineeringReference-Based Inputs

Qualifications

5+ Years Product Evaluation Experience5+ Years UX Research Experience5+ Years Model Testing Experience5+ Years Structured Qualitative Assessment ExperienceMaster's or Higher in Cognitive ScienceMaster's or Higher in Design ResearchMaster's or Higher in Media StudiesMaster's or Higher in Related Field

Full job description

You'll own how Luma judges whether its models are actually good, past the point where numbers stop telling the story. As our Qualitative Evaluation Engineer, you'll build the frameworks that pin down believability, identity retention, and scene coherence, and turn them into insight that steers model development.

This isn't a checkbox-metrics role. You're building evaluative systems that match the messiness of human perception and creative intent, and much of that framework doesn't exist yet. It fits someone who can take a fuzzy quality and define it in clear, testable terms, working shoulder to shoulder with researchers and technical artists. If you want work scored purely on quantitative dashboards, this isn't it.

What You'll Own

  • Evaluate generative model performance across diverse tasks, prompts, and modalities, and surface the failure modes, regressions, and edge cases that hurt product quality.

  • Build and maintain qualitative evaluation frameworks that are scalable and reusable.

  • Translate high-level product goals into concrete evaluative criteria.

  • Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluations.

  • Turn nuanced judgments into clear feedback that informs fine-tuning, dataset curation, and product UX.

  • Work closely with technical artists and engineers to keep evaluations aligned with model capabilities and real use cases.

First 90 Days

One way the first 90 could unfold.

  • Days 1–30 — Immerse & Diagnose: Learn the models and the creative use cases they serve, and audit how quality is judged today and where it misses.

  • Days 30–60 — Ship & Validate: Stand up a qualitative framework for one high-priority capability and run it on real outputs, producing insight the team acts on.

  • Days 60–90 — Scale & Systemize: Make the framework reusable across capabilities and wire it into the model/data/eval loop.

What You Bring

  • 5+ years in product evaluation, UX research, model testing, or similar structured qualitative assessment.

  • Master's or higher in Cognitive Science, HCI, Design Research, Psychology, Media Studies, or a related field.

  • Deep familiarity with creative workflows for generative models (animation, filmmaking, digital art, VFX).

  • Systems thinking: you can define abstract qualities like believability or scene coherence in clear evaluative terms.

  • Excellent written communication and the ability to synthesize nuanced judgment into actionable insight.

  • Comfort working across engineers, researchers, and creatives.

Nice to Have

  • Background in motion, visual effects, or storytelling pipelines.

  • Experience evaluating AI-generated media (video, images, 3D).

  • Prior work building internal tools for qualitative data collection or scoring.

  • Familiarity with prompt engineering and reference-based inputs.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

You've read the whole posting — now see how you match it.