Staff Research Engineer, Agent Evals & Post-training
San Francisco, CA, USAFull-timePosted 4 days agoStill listed 3 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
Nominal is hiring a Staff Research Engineer to define how its agents are evaluated and to build a path from evaluation to post-training when needed. The role focuses first on measuring agent performance on real hardware tasks, then on training models for needs such as air-gapped deployment or cost. The engineer will set measurement and model-improvement strategy with hardware experts.
Skills & qualifications
Skills
Qualifications
Benefits
Full job description
Staff Research Engineer, Agent Evals & Post-training New York, NY • Los Angeles, CA • London • Austin, TX • San Francisco, CA • Seattle, WA
AI
In office
Full-time
About Nominal
Our mission is to accelerate how the world engineers new hardware. Nominal's connected test and operations platform powers the world's most advanced hardware programs and its most ambitious startups, from spacecraft, racecars, and autonomous vehicles to next-generation defense and energy programs. Our customers include Anduril, Shield AI, Hermeus, Albedo, Shinkei, and Pratt Miller Motorsports, as well as U.S. Navy and U.S. Air Force programs. Now we're expanding across the entire hardware lifecycle, building the foundation, AI-native applications, and agents that accelerate innovators' work and change what's possible to build.
We're backed by Sequoia, General Catalyst, Founders Fund, Lux Capital, and Lightspeed, and our team comes from SpaceX, Apple, Palantir, Anduril, Applied Intuition, and other leading companies.
About Hardware Intelligence
We're the team behind Nominal's agents, AI-native applications, and MCP, and its forward-leaning AI bets. Our mission is to unlock the bottlenecks of the hardware lifecycle with AI. Our agents reason over physical reality, from high-rate telemetry and test campaigns to designs and simulations, where real test results are the ground truth their work is checked against. We believe opinionated AI, built for the real work of hardware programs, will change how the world engineers.
We're collaborative, iterative, and high-agency, and we're human-centered and customer-focused. We build with the newest AI tools every day, and because those tools keep changing, so do we: we stay curious and keep looking for the better way. Our team spans data science and ML, distributed systems, search, and knowledge systems, and we obsess over how agents can be genuinely useful to the engineers who rely on them.
The Role:
As a Staff Research Engineer supporting agent evals & post-training, you'll define how Nominal measures its agents, for customers and for ourselves, and build the path from evals to post-trained models when hardware needs them. Evals come first; post-training follows when air-gapped deployment or cost makes it the right investment.
💼 What You'll Do:
- Build eval suites for our agents, the MCP tool layer, and our internal company agent, grounded in real hardware tasks.
- Invent new benchmarks for agents working over physical engineering data, where none exist today.
- Benchmark our agents against frontier agents using our MCP, and show where and why ours win.
- Build the eval infrastructure: datasets, model-based and human graders, and regression gates in CI.
- Turn evals into reward signals and training data, and lead post-training (fine-tuning, RL, distillation) when we need models we can run anywhere, including air-gapped environments.
- Set Nominal's strategy for measurement and model improvement, working with hardware experts on what "correct" means.
🚀 What you'll bring:
- 8+ years in ML engineering or research, including evals or post-training work you led in production.
- Statistical rigor: you design evals that don't fool you, and you know when a difference is real.
- Deep experience evaluating LLM or agent systems, including model-graded evals and their limits.
- Hands-on post-training experience: fine-tuning, RL from feedback, or distillation on real tasks.
- The judgment to know when to measure, when to train, and when a better prompt or tool is the answer.
- A track record of setting technical direction across a team and raising the bar for the engineers around you.
- You build with modern AI coding agents (Claude Code, Cursor, Codex) every day, and stay curious and open to better ways of working. The tools keep changing, and so do we.
⚡️ Nice to have:
- You've built evals or post-training at a frontier lab or an AI-native company.
- You've published benchmarks or eval methods that others use.
- You've trained or served open-weight models in restricted, on-prem, or air-gapped environments.
- You've worked in test, reliability, or verification engineering for physical systems.
Benefits/Perks
- 🏥 100% coverage of medical, dental, and vision insurance
- 🏖️ Unlimited PTO and sick leave
- 🍽️ Free lunch, snacks, and coffee
- 🚀 Professional Development Stipend
- ✈️ Annual company retreat
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin.
ITAR Requirements To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.
Ready to apply? Powered by
First name *
Last name *
Email *
LinkedIn URL *
Phone number *
Location *
Resume * Click to upload or drag and drop here
Cover letter Click to upload or drag and drop here
This position requires access to information and technology that is subject to U.S. export controls. Your responses to the questions below will be used solely to determine your eligibility under U.S. law to receive information and materials subject to U.S. export controls. Are you any of the following “protected individual(s)” as defined in the Immigration and Naturalization Act, 8 U.S.C. 1324b(a)(3)? *
A person lawfully admitted for permanent residence of the United States (i.e. Green Card holder)
A person admitted as a refugee to the United States under 8 U.S.C. 1157
A person admitted as an asylee to the United States under 8 U.S.C. 1158
A United States citizen or national
None of the above
U.S. Work Authorization - Are you authorized to work in the United States? *
Yes
No
Will you require sponsorship from Nominal for employment now or in the future? (e.g, OPT, H1B visa)? *
Yes
No
Clearance Eligibility - Do you presently hold an active U.S. security clearance, or are you eligible to obtain and maintain a U.S. security clearance? *
Yes, I hold an active U.S. security clearance
Yes, I'm eligible for a U.S security clearance
No
If you have held a U.S. security clearance in the past, what clearance level have you held? * Confidential
Secret
Top Secret
N/A - have never held a U.S. security clearance
Are you open to working 5 days onsite in one of our offices listed on the job? *
Yes
No
Voluntary Self-Identification To comply with government reporting requirements, we invite candidates to participate in the self-identification survey below. Your completion of this form is entirely optional, and your decision will neither influence the hiring process nor any subsequent stages. Any information you choose to share will be kept confidential and stored in a secure file. As outlined in our Equal Employment Opportunity policy, we uphold a commitment to non-discrimination based on any protected group status specified in applicable laws. Gender
Race
Race and ethnicity descriptions
Voluntary Self-Identification of Veteran Status VEVRAA requires Government contractors to take affirmative action to employ and advance in employment protected veterans. To help us measure the effectiveness of our outreach and recruitment efforts of veterans, we are asking you to tell us if you are a veteran covered by VEVRAA. If you believe that you belong to any of the following categories of protected veterans, please indicate by making the appropriate selection. Veteran status descriptions Disabled veteran A veteran who served on active duty in the U.S. military and is entitled to disability compensation (or who but for the receipt of military retired pay would be entitled to disability compensation) under laws administered by the Secretary of Veterans Affairs, or was discharged or released from active duty because of a service-connected disability Recently separated veteran A veteran separated during the three-year period beginning on the date of the veteran's discharge or release from active duty in the U.S military, ground, naval, or air service Active duty wartime or campaign badge veteran A veteran who served on active duty in the U.S. military during a war, or in a campaign or expedition for which a campaign badge was authorized under the laws administered by the Department of Defense Armed Forces service medal veteran A veteran who, while serving on active duty in the U.S. military ground, naval, or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985 (61 Fed. Reg. 1209).
Veteran status
Req ID: R146
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
2027 Summer Intern, PhD, Research, Post TrainingWaymo · Mountain View, CA (Hybrid) · $85/hrPosted todayPosted today
Research Scientist, Post-TrainingDatologyAI · San Mateo, CA (Hybrid) · $180–300K/yrPosted 2 days agoPosted 2 days agoPost-Training Applied Researcherbaseten · San Francisco, CA (Hybrid) · $200–275K/yrPosted 3 days agoPosted 3 days ago
Member of Technical Staff, EvalsHandshake · San Francisco, CA · $200–350K/yrPosted 1w agoPosted 1w ago
Research Engineer, Robotics Evalshud · San Francisco, CA · $100–230K/yrPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.