Sage Care logo

Software Engineer, Agent Intelligence

Sage Care

Marquette, MIHybridFull-time$200–250K/yrPosted 5mo agoVerified open 5 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$200–250K/yr
Location
Marquette, MIHybrid
Schedule
Full-time
Work Authorization
Not specified

Job overview

Sage Care is hiring a Software Engineer, Agent Intelligence. Sage Care seeks a Software Engineer to own the intelligence layer of its AI‑agent platform, building evaluation pipelines, feedback systems, and ML infrastructure that turn real patient conversations into continuous improvement. The role bridges AI evaluation, production quality, and operations to make agent failures diagnosable and automate learning loops.

Key focus areas include Design systems that continuously analyze production conversations and identify quality issues, Build evaluation pipelines that measure agent performance across patient‑care dimensions, and Develop workflows that transform human feedback into actionable improvements.

Successful candidates bring 5+ Years Of Software Engineering Experience. Important skills include Building Production Systems, Working With LLMs, Working With AI Agents, Conversational AI Applications, Building AI Evaluation Platforms, and Building AI Evaluation Frameworks. Preferred (not required): Model Quality Measurement, Voice AI Systems, Designing Human-In-The-Loop Workflows, and Turning Emerging Research Into Tested Prototypes.

Skills & qualifications

RequiredNice to have

Skills

Building Production SystemsWorking With LLMsWorking With AI AgentsConversational AI ApplicationsBuilding AI Evaluation PlatformsBuilding AI Evaluation FrameworksDeveloping Feedback Systems for ML ApplicationsCreating Observability Tooling for AI ProductsBuilding ML Pipelines to Improve Model PerformanceBackend EngineeringSystems ThinkingWorking With Ambiguous ProblemsDefining Solutions From First PrinciplesModel Quality MeasurementVoice AI SystemsDesigning Human-in-the-Loop WorkflowsTurning Emerging Research Into Tested PrototypesTurning New Techniques Into Tested PrototypesBuilding Internal PlatformsHealthcare Domain ExperienceRegulated Domain ExperienceSafety-Critical Domain ExperienceBuilding AI Evaluation Platforms or FrameworksDeveloping Feedback Systems for ML and LLM ApplicationsBuilding ML Pipelines That Improve Model or Agent PerformanceHealthcare DomainRegulated DomainsSafety-Critical Domains

Qualifications

5+ Years of Software Engineering Experience

Full job description

ABOUT SAGE CARE

Sage Care is a fast-growing, early-stage healthcare startup founded by exceptional leaders from Apple, Uber, Carbon Health and backed by top-tier venture capital (General Catalyst, Chelsea Clinton). With a strong customer pipeline, Sage Care is transforming healthcare by simplifying care navigation.

Our platform makes it easier for patients to find the right doctor and helps providers focus on those who need them most through harnessing the latest AI innovations.

Building on our successful collaborations with health systems across the U.S., we have expanded internationally to the MENA region. We are now partnering with health systems there to deploy our AI-powered care navigation platform.

ABOUT THE ROLE

Every day, our AI agents handle real patient conversations for hospital systems. Those conversations contain everything needed to make the agents better. Today, much of that learning is still manual: humans review calls, identify issues, investigate failures, and work with engineers to improve agent behavior.

Your mission is to build the systems that close that loop.

You will own the intelligence layer of our agent platform: the evaluation pipelines, feedback systems, and ML infrastructure that turn production conversations into measurable, continuous improvement.

This role sits at the intersection of AI evaluation, ML pipelines, and production quality. You will work closely with the engineers who own the agent runtime, with the operations teams who review calls, and with product leaders who decide what good looks like for each hospital partner.

WHAT YOU'LL DO

Build the evaluation and feedback platform

  • Design systems that continuously analyze production conversations, identify and cluster quality issues

  • Build evaluation pipelines that measure agent performance across the dimensions that matter for patient care

  • Develop workflows that transform human feedback into actionable improvements

  • Design mechanisms for routing issues to the appropriate AI, engineering, or operational owners

Make failures diagnosable

  • Investigate production failures and identify root causes across transcription, reasoning, retrieval, and orchestration

  • Build tooling that helps engineers quickly understand why an agent behaved a certain way

  • Establish quality metrics and reliability standards for production agents

Automate the learning loop

  • Build ML pipelines that reduce the manual effort required to improve agents

WHAT WE'RE LOOKING FOR

Required

  • 5+ years of software engineering experience

  • Experience building production systems

  • Experience working with LLMs, AI agents, or conversational AI applications

  • Experience in one or more of the following:

    • building AI evaluation platforms or frameworks

    • developing feedback systems for ML and LLM applications

    • creating observability, reliability, or quality tooling for AI products

    • building ML pipelines that improve model or agent performance

  • Strong backend engineering skills and systems thinking

  • Experience working with ambiguous problems and defining solutions from first principles

Nice to Have

  • Experience with evaluation frameworks and model quality measurement at AI-native companies

  • Experience with voice AI systems

  • Experience designing human-in-the-loop workflows

  • Experience turning emerging research or new techniques into tested prototypes

  • Experience building internal platforms used by engineering teams

  • Experience in healthcare or other regulated, safety-critical domains

WHAT SUCCESS LOOKS LIKE

  • Failures are detected by systems, not discovered by people

  • Every production issue has a diagnosable root cause and a clear owner

  • Human feedback measurably changes agent behavior, with the lag between the two shrinking over time

You've read the whole posting — now see how you match it.