Gap Inc (Gap/Old Navy/Athleta) logo

Principal Engineer, Agent Platform

Gap Inc (Gap/Old Navy/Athleta)

Remote · USFull-timePosted 1 day agoStill listed today

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Remote · US
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Gap Inc. is building an enterprise AI agent platform for teams to create, deploy, govern, and operate AI-powered solutions at scale. The Principal Engineer will define its technical foundation and build platform services, frameworks, and guardrails. This hands-on role spans architecture, software design, and platform development across AI, distributed systems, cloud infrastructure, security, and governance.

Skills & qualifications

RequiredNice to have

Skills

Distributed SystemsPythonLLM SolutionsGenerative AIAgent ArchitecturesOrchestration PatternsEnterprise AI PlatformsRetrieval SystemsEmbeddingsVector SearchContext ManagementAI EvaluationTesting MethodologiesQuality AssuranceGoogle CloudVertex AIGeminiGKECloud RunBigQueryPub/SubCloud SecurityWorkload IdentityOAuthPolicy-as-CodeSecure ArchitectureReliability PatternsResiliencyRetriesObservabilityFault ToleranceTechnical InfluenceArchitectural StandardsTechnical Leadership

Qualifications

Extensive Production Experience

Full job description

About the Role We are building the enterprise AI agent platform that teams across the company will use to create, deploy, govern, and operate AI-powered solutions at scale. As a Principal Engineer, you will help define the technical foundation for agent development across the organization, building the services, frameworks, and guardrails that make AI agents secure, reliable, observable, and easy to adopt.

This is a hands-on engineering role focused on architecture, software design, and platform development. You will work across AI, distributed systems, cloud infrastructure, security, and governance to deliver a platform that enables product and engineering teams to move quickly while meeting enterprise standards for security, compliance, and operational excellence.What You'll Do

  • Design and build core platform services that enable enterprise AI agent development and operations

  • Define the architecture, engineering patterns, and technical standards for the AI agent platform

  • Develop and maintain a centralized model gateway for routing, governance, resiliency, and cost management

  • Build agent runtime capabilities that support long-running workflows, orchestration, and human-in-the-loop processes

  • Design secure identity, authorization, and policy enforcement mechanisms for agents and tools

  • Establish platform capabilities for context management, memory, retrieval, and data isolation

  • Create testing, evaluation, and release frameworks that ensure agent quality, safety, and performance

  • Implement observability, auditability, and operational tooling for monitoring and troubleshooting agent behavior

  • Partner with Security, Data Governance, and Cloud Infrastructure teams to define platform controls and guardrails

  • Mentor engineers, influence technical direction, and raise engineering standards across multiple teams

Who You Are

  • Extensive experience building and operating large-scale distributed systems in production

  • Strong software engineering background with expert-level Python development skills

  • Hands-on experience delivering production LLM, generative AI, or agent-based solutions

  • Deep understanding of agent architectures, orchestration patterns, and enterprise AI platforms

  • Experienced with retrieval systems, embeddings, vector search, and context management strategies

  • Knowledgeable in AI evaluation, testing methodologies, and quality assurance for non-deterministic systems

  • Proficient with Google Cloud services including Vertex AI, Gemini, GKE, Cloud Run, BigQuery, and Pub/Sub

  • Experienced with cloud security, workload identity, OAuth, policy-as-code, and secure-by-design architectures

  • Strong understanding of distributed system reliability patterns, including resiliency, retries, observability, and fault tolerance

  • Proven ability to influence technical direction, drive adoption of architectural standards, and lead through expertise rather than authority

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.