Sage Care logo

Software Engineer, Agent Intelligence

Sage Care

Palo Alto, CAJobNo compensation foundPosted 7mo agoVerified open 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Palo Alto, CA
Work Authorization
Not specified

Job overview

Sage Care is hiring a Software Engineer, Agent Intelligence. Sage Care, an early‑stage healthcare startup, is building AI‑powered care navigation platforms that help patients find doctors and enable providers to focus on those who need them most. The role focuses on creating evaluation pipelines, feedback systems, and ML infrastructure to turn production conversations into continuous improvement for AI agents.

Key focus areas include Design systems that continuously analyze production conversations, identify and cluster quality issues, Build evaluation pipelines that measure agent performance across the dimensions that matter for patient care, and Develop workflows that transform human feedback into actionable improvements.

Successful candidates bring 5+ Years Software Engineering Experience. Important skills include Python, Java, Rust, LLMs, AI Agents, and Conversational AI Applications. Preferred (not required): Building AI Evaluation Platforms, Developing Feedback Systems For ML And LLM Applications, Conversational AI Systems, and Creating Observability, Reliability, Or Quality Tooling For AI Products.

Skills & qualifications

RequiredNice to have

Skills

PythonJavaRustLLMsAI AgentsConversational AI ApplicationsBackend EngineeringSystems ThinkingDefining Solutions From First PrinciplesCommunicationCollaborationBuilding AI Evaluation PlatformsDeveloping Feedback Systems for ML and LLM ApplicationsConversational AI SystemsCreating Observability, Reliability, or Quality Tooling for AI ProductsDesigning Automated Workflows That Improve Model or Agent PerformanceBuilding Production AI ApplicationsEvaluation FrameworksModel Quality MeasurementVoice AI SystemsDesigning Human-in-the-Loop WorkflowsBuilding Internal PlatformsAI Evaluation PlatformsFeedback Systems for ML and LLM ApplicationsObservabilityReliabilityQuality Tooling for AI ProductsML PipelinesAI Evaluation FrameworksHuman-in-the-Loop WorkflowsTurning Emerging Research Into Tested PrototypesHealthcare DomainRegulated DomainsSafety-Critical Domains

Qualifications

5+ Years of Software Engineering ExperienceExperience Building Production SystemsExperience Working With Ambiguous Problems

Full job description

About Sage Care Sage Care is a fast-growing, early-stage healthcare startup founded by exceptional leaders from Apple, Uber, Carbon Health and backed by top-tier venture capital (General Catalyst, Chelsea Clinton). With a strong customer pipeline, Sage Care is transforming healthcare by simplifying care navigation.

Our platform makes it easier for patients to find the right doctor and helps providers focus on those who need them most through harnessing the latest AI innovations.

Building on our successful collaborations with health systems across the U.S., we have expanded internationally to the MENA region. We are now partnering with health systems there to deploy our AI-powered care navigation platform.

About the Role Every day, our AI agents handle real patient conversations for hospital systems. Those conversations contain everything needed to make the agents better. Today, much of that learning is still manual: humans review calls, identify issues, investigate failures, and work with engineers to improve agent behavior.

Your mission is to build the systems that close that loop.

You will own the intelligence layer of our agent platform: the evaluation pipelines, feedback systems, and ML infrastructure that turn production conversations into measurable, continuous improvement.

This role sits at the intersection of AI evaluation, ML pipelines, and production quality. You will work closely with the engineers who own the agent runtime, with the operations teams who review calls, and with product leaders who decide what good looks like for each hospital partner.

What You'll Do Build the evaluation and feedback platform

  • Design systems that continuously analyze production conversations, identify and cluster quality issues

  • Build evaluation pipelines that measure agent performance across the dimensions that matter for patient care

  • Develop workflows that transform human feedback into actionable improvements

  • Design mechanisms for routing issues to the appropriate AI, engineering, or operational owners

Make failures diagnosable

  • Investigate production failures and identify root causes across transcription, reasoning, retrieval, and orchestration

  • Build tooling that helps engineers quickly understand why an agent behaved a certain way

  • Establish quality metrics and reliability standards for production agents

Automate the learning loop

  • Build ML pipelines that reduce the manual effort required to improve agents

What We're Looking For Required

  • 5+ years of software engineering experience

  • Experience building production systems

  • Experience working with LLMs, AI agents, or conversational AI applications

  • Experience in one or more of the following:

  • building AI evaluation platforms or frameworks

  • developing feedback systems for ML and LLM applications

  • creating observability, reliability, or quality tooling for AI products

  • building ML pipelines that improve model or agent performance

  • Strong backend engineering skills and systems thinking

  • Experience working with ambiguous problems and defining solutions from first principles

Nice to Have

  • Experience with evaluation frameworks and model quality measurement at AI-native companies

  • Experience with voice AI systems

  • Experience designing human-in-the-loop workflows

  • Experience turning emerging research or new techniques into tested prototypes

  • Experience building internal platforms used by engineering teams

  • Experience in healthcare or other regulated, safety-critical domains

What Success Looks Like

  • Failures are detected by systems, not discovered by people

  • Every production issue has a diagnosable root cause and a clear owner

  • Human feedback measurably changes agent behavior, with the lag between the two shrinking over time

You've read the whole posting — now see how you match it.