
Applied AI Engineer, Agent Harness
Cupertino, CAFull-timeSeen todaySeen in employer's feed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Cupertino, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
The Applied AI Engineer will build the core runtime harness that enables Siri’s models to act as an assistant, handling context, tool calls, streaming, and multi-step tasks across Apple devices and Private Cloud Compute. The role focuses on real-time systems under latency, memory, and power constraints, using evaluations to understand and improve model behavior. The engineer will collaborate across modeling, tools, context, app, and framework teams and take features from design through production.
Skills & qualifications
Skills
Qualifications
Full job description
Weekly Hours: 40
Role Number: 200688056-0836
Summary
Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world’s most widely used AI assistants. It runs on our next generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, and it is built with privacy from the ground up. We build the harness, the foundation that turns a model into an assistant that can act. This is a systems engineering role for people who understand how large language models behave.
Description
A model can't act on its own. The harness runs the agent loop, carries context and tools to the model, executes the tool calls it makes, and streams results back reliably. It spans models running on device and larger models in Private Cloud Compute. It handles everything from single-turn requests to long multi-step tasks that span apps, survive interruptions, and resume cleanly.
We're looking for a strong systems engineer who is fluent in how LLMs behave. The work is real-time and resource-constrained: responses must start fast enough for a voice assistant to feel instant, under tight latency, memory, and power limits on everything from Mac and iPhone to Apple Watch. Failures rarely have a single cause. In a typical week you might track down why a streamed response stalls, make tool calls cancel cleanly when the user interrupts, or replay a failed request to tell whether the model, the context, or the runtime was at fault. Every change to the harness changes how the model behaves, so we measure our work with evals, not just tests.
You'll join the team that owns the core runtime and work alongside senior engineers and researchers. You'll also partner closely with the modeling, tools, and context teams and with app and framework teams across Apple. You'll take features from design through production and grow your ownership of the system over time. You'll use agent harnesses every day while building one for an assistant used by hundreds of millions of people.
Minimum Qualifications
-
Bachelor's degree in Computer Science or a related field, with 2+ years of industry experience
-
Strong programming skills in a systems language such as Swift, C++, Rust, Objective-C, or Go
-
Solid foundation in concurrency, inter-process communication, and performance optimization
-
Hands-on experience building LLM-powered systems such as agents, tool calling, or prompt and context design, through professional work or substantial personal projects
-
Daily use of agentic coding tools such as Claude Code, Codex, or Pi, with informed opinions on what makes their harnesses succeed or fail
-
Ability to communicate clearly and make progress on ambiguous problems
Preferred Qualifications
-
Experience with Swift and Apple platform development
-
Experience building agent harnesses, orchestration frameworks, or multi-step tool-use systems
-
Experience with model inference integration, streaming protocols, or latency-sensitive client-server systems
-
Experience using evals to diagnose and improve LLM behavior
-
Experience designing APIs or platforms adopted by other engineering teams
-
Familiarity with privacy and security engineering in production systems
-
Active engagement with current research and practice in agent harness design, such as context engineering, tool use, and long-horizon agents, with a record of turning new ideas into practical improvements
Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Harness Design Intern (Summer 2027)Muon Space · San Jose, CA · $40–50/hrPosted 1w agoPosted 1w ago
Senior Product Manager, Agent OptimizationNVIDIA · Remote · US · $208–328K/yrPosted 1w agoPosted 1w agoMTS, Agent ResearchEtched · San Jose, CA · $175–275K/yrPosted 2w agoPosted 2w ago
Senior Applied AI EngineerNVIDIA · Santa Clara, CA · $152–242K/yrPosted 2 days agoPosted 2 days agoApplied Engineer, HardwareEtched · San Jose, CA · $150–250K/yrPosted 2 days agoPosted 2 days ago
You've read the whole posting — now see how you match it.