Evals Lead
Washington, DCFull-timePosted 2w agoStill listed 1 day ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Washington, DC, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
Aslan seeks an Evals Lead to operate and improve autonomous agent platforms, building tooling, datasets, and regression suites to assess agent decisions. The role involves daily platform monitoring, diagnosing model versus harness issues, defining regression metrics, and creating simulation harnesses, all within a national security context.
Skills & qualifications
Skills
Qualifications
Full job description
About Aslan Our national security apparatus was designed for an adversary that congregated and operate in the physical world. The threats disrupting our way of life today have gone faceless, borderless, and beyond the reach of any human operator. Aslan builds the autonomous agents that reach them, unmask them, and stop them, before a life can be lost, a household can be bankrupted, innocence can be stolen, or our can be nation subverted. In 12 months we've gone from founding to live operational tasking and pilots across multiple national security and law enforcement partners. We've proven it works. This role makes it permanent. We’ve raised $20M to date from Khosla Ventures, XYZ, BoxGroup, 2048, Liquid 2, Precursor Ventures, and others.
Role Aslan builds autonomous agents that run for weeks at a time across many external systems, plan their own next steps, and work toward an objective. Scoring them the usual way doesn't work. Individual outputs look fine. Results come back late or not at all. What we need to know is whether the agent chose well at the moment it chose, and nothing off the shelf measures that. This is our first quality role. You'll operate the platform daily, build the tooling that catches what you catch by hand, and own the decision to ship. We deploy on-premise on a fixed cadence and can't push fixes after delivery.
Responsibilities
-
Run the platform yourself every day and log what's wrong in enough detail that an engineer can reproduce it.
-
Sample agent plans, get them rated on whether the agent made the right call, and turn the ratings into a versioned dataset. Engineering uses it as a regression suite. The training side uses it as labels.
-
Diagnose whether a bad decision came from the model or from the harness handing it the wrong context, and route each to the right owner.
-
Write and run checks over each agent's full history, and over properties that only hold across the whole set of running agents.
-
Build a simulation harness so a week of agent operation runs overnight against a candidate build.
-
Define what counts as a regression between builds. Measure both whether agents last and whether they get anywhere. A build that only makes them cautious is a failed build.
Who You Are
-
You've owned evals for a shipped LLM or agent product, and you can say what your suite caught and what it missed
-
You've built a labeled dataset that got used for training, not just for measurement
-
You've picked a proxy metric because the real signal wasn't available, and you can defend the choice
-
You know when LLM-as-judge works and when it doesn't
-
You'd rather write the check than file the ticket
What You Bring
-
Strong Python
-
An evals framework you've run in production: Inspect, Promptfoo, Braintrust, Arize, LangSmith or equivalent
-
Familiarity with agent harness internals: memory, context assembly, tool selection
-
Preference data, process supervision, or reward modeling experience
-
Enough statistics to tell a regression from noise on small samples
Bonus Points
-
Evaluating agents that run for days or weeks, where per-episode metrics don't work
-
Systems you can't roll back or reset
-
Graph stores, provenance, lineage tracking
-
Eligible for a US security clearance
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Cyber Enablement Lead, GovOpenAI · Washington, DC (Hybrid) · $207–230K/yrPosted 2 days agoPosted 2 days ago
Security Data, Analytics, and Technology LeadFresenius Medical Care · Remote · USPosted 1w agoPosted 1w ago
Lead - US Government Affairs & Public PolicyCohere · Washington, DC, DC (Hybrid) · $175–330K/yrPosted 4w agoPosted 4w agoLead Director, Risk Adjustment Performance Reporting and InsightsCVS Health · Remote · US · $100–232K/yrPosted 1w agoPosted 1w ago
Capture Strategy & Operations Sr. Lead, X-BAT Family of Systems (FoS) (R6083)Shield AI · San Diego, CA · $160–300K/yrPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.