Member of Technical Staff - Agent Engineer
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
Phylo is hiring a Member of Technical Staff - Agent Engineer. Phylo seeks an engineer with research or production experience in AI agents to build, evaluate, and improve agent harnesses, advancing multi‑agent coordination, model routing, memory, planning, and tool use for scientific discovery, while owning production systems and collaborating with scientists.
Key focus areas include Advance the agent harness by integrating latest research and open‑source developments into production and experimenting with multi‑agent coordination, model routing, memory, planning, and tool use., Build rigorous evaluations measuring agent quality, reliability, latency, and cost on scientific tasks., and Monitor agent quality in production, troubleshoot failures, and translate findings into engineering improvements..
Successful candidates bring Research Or Production Experience In AI Agents, Research Or Production Experience In Machine Learning Or Quantitative Discipline, and Experience Working In AI-Native Teams. Important skills include Software Engineering, Backend Systems, Distributed Systems, Machine Learning, LLM APIs, and Tool Calling. Preferred (not required): LLM Evaluations, Human Evaluation, Model Judges, and Replay Testing.
Skills & qualifications
Skills
Qualifications
Benefits
Full job description
Member of Technical Staff – Agent Engineer
About Phylo
Phylo is an applied research lab building agentic intelligence to accelerate discovery for every biomedical scientist. We believe AI agents will fundamentally transform how biomedical research is done. Our team brings together engineers, AI researchers, and scientists to build systems capable of carrying out complex scientific work.
About the role
We’re looking for an engineer with research or production experience in AI agents to build and evaluate systems that make our agents capable, reliable, and measurably better.
The model is only one part of an agent; the harness shapes how it plans, uses tools, manages context, coordinates work, and recovers from failure. You will own these systems in production and build evaluations to measure whether changes improve performance. The work includes bringing the latest advances and original ideas into production. The ideal candidate combines strong engineering execution with a quantitative mindset and an interest in AI agents for scientific discovery.
What you’ll work on
-
Advance the agent harness by bringing the latest research and open-source developments into production and experimenting with new approaches to multi-agent coordination, model routing, memory, planning, and tool use.
-
Build rigorous evaluations that measure agent quality, reliability, latency, and cost on representative scientific tasks.
-
Monitor agent quality in production, troubleshoot failures quickly in high-ambiguity situations, and translate findings into engineering improvements.
-
Collaborate on reliable infrastructure for long-running sessions, safe sandboxed execution, background work, and subagents.
-
Partner with scientists and engineers to evaluate and productionize new agent capabilities.
Requirements
-
Strong software engineering experience building production backend systems, infrastructure, or distributed systems.
-
Research or production experience in machine learning, or another quantitative discipline.
-
Familiarity with LLM APIs, tool calling, agent runtimes, or workflow orchestration.
-
Strong quantitative judgment and the ability to determine whether an apparent improvement is real, reproducible, and meaningful.
-
Ability to move between research questions, data analysis, system design, and production implementation.
-
Experience or strong interest in AI for science, scientific agents, computational research, or automated scientific discovery.
-
Experience working in AI-native teams that use coding agents or automation extensively.
-
High ownership, clear communication, and strong engineering judgment.
We care more about demonstrated ability than a particular credential. Relevant backgrounds may include ML or LLM engineering, academic research paired with substantial software development, scientific computing, or production systems engineering with demonstrated quantitative experience.
Nice to have
-
Experience with LLM evaluations, human evaluation, model judges, replay testing, benchmark design, or experiment tracking.
-
Experience applying classical machine learning methods alongside LLMs in production systems.
-
Experience with task queues, event streams, Kubernetes, code sandboxes, or durable workflow systems.
-
An advanced degree or equivalent experience in machine learning, computational biology, physics, applied mathematics, statistics, or a related field.
Why Join Us?
-
Competitive salary and equity share in building the future of biomedical discovery
-
Full medical, dental, and vision coverage, including free therapy sessions and eyewear stipend
-
401(k) to help you build long-term financial security (US only)
-
Unlimited PTO to recharge when you need it (US only)
-
Lunch and snacks when you're in the office
-
Regular team offsites and company events
-
A culture of excellence and speed - we move fast, think big, and support each other every step of the way
-
Your work will directly impact our mission to 100X biomedical discoveries through AI
You've read the whole posting — now see how you match it.