Zoox logo

Machine Learning Automation Engineer

Zoox

Foster City, CAHybridFull-time$160–225K/yrPosted 1mo agoVerified open 6 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$160–225K/yr
Location
Foster City, CAHybrid
Schedule
Full-time
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Master's degree

Job overview

The role focuses on building and operating ML, deep learning, and LLM‑powered triage systems within Zoox’s System Behavior Analysis organization, creating data pipelines that detect and resolve robot failures, and collaborating with safety, planning, and quality assurance stakeholders to keep production pipelines healthy.

Skills & qualifications

RequiredNice to have

Skills

Machine LearningDeep LearningAgentic LLM SystemsPythonCI/CDDashboardsData ManipulationRAGPyTorchTensorFlowHuggingFaceDatabricksAWSSelf Correcting ModelsInstallation and Maintenance of LLMsData Visualization ToolsLookerDashPerformance Monitoring at ScaleSupport EnvironmentsAutonomous DrivingRoboticsVerbal CommunicationWritten Communication

Qualifications

PhD or Master's in STEM Field4+ Years in Triage, Performance Analytics, Infrastructure as Code, ML, NLP, or Deep Learning

Full job description

Within the System Behavior Analysis (SBA) organization, the Behavioral Automated Triage (BAT) team builds the ML, deep learning, and data pipelines that help Zoox understand and triage failures our robots encounter on the road. You will build agentic, LLM-powered triage systems and run them in production, keeping pipelines healthy that teams across Zoox use daily, while working closely with stakeholders across safety, planning, and quality assurance.

In this role, you will:

  • Design and improve ML/DL algorithms on large-scale data to automate test and triage workflows.

  • Build and run agentic LLM systems (e.g., Claude or Gemini) that automate triage, from prototype to production.

  • Run multiple production pipelines: monitor them, respond to issues, and fix the root causes.

  • Work with stakeholders across data science, ML, autonomous drive planning, and quality assurance.

  • Build monitoring tools and use them to keep improving your algorithms in production.

Qualifications:

  • PhD or Master's in a STEM field and 4+ years in Triage, Performance Analytics, for fortune 500 including infrastructure as code, ML, NLP, or deep learning.

  • Strong productionization, CI/CD experience for medium to large projects Python,, with a solid grasp of dashboards, eval, algorithms and data manipulation.

  • Experience with RAG, PyTorch, TensorFlow, or HuggingFace

  • Has built or managed agentic LLM systems (Claude or Gemini) in a professional setting with large dataset to satisfy multiple stakeholders.

  • Has run multiple production pipelines on platforms like Databricks and AWS, including fixing production issues.

Bonus Qualifications:

  • Experience with self correcting models, installation and maintenance of LLMs, data visualization tools like Looker/Dash/Databricks Dashboards

  • Experience in Performance Monitoring at Scale, Support environments, autonomous driving, robotics.

  • Strong verbal and written communication skills

You've read the whole posting — now see how you match it.