Apple logo

Software Engineer - Generative Data Platform, Evaluation

Apple

San Francisco, CAFull-timeSeen 1 day agoSeen in employer's feed 1 day ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Apple's AI & Machine Learning team builds a self‑service platform that generates synthetic datasets for privacy‑preserving training and evaluation, helping engineers, data scientists, designers and safety teams create realistic data at scale.

Skills & qualifications

RequiredNice to have

Skills

PythonDistributed SystemsData PipelinesGPU InfrastructureInput ValidationConfiguration ContractsMachine Learning ConceptsAI Coding ToolsGenerative Image/Video APIsStructured OutputImage and Video File FormatsMetadata StandardsBatch Compute PlatformsJob Schedulers

Qualifications

Bachelor's Degree in Computer Science or Related Field3+ Years Building Production Software in PythonExperience Debugging Distributed Systems or Data PipelinesExperience Deploying Machine Learning Models on GPU InfrastructureExperience Designing Input Validation and Configuration ContractsAbility to Explain Machine Learning Concepts to Non‑Technical PartnersHands‑on Use of Agentic AI Coding ToolsExperience Working Directly With Partner Teams to Define and Ship FeaturesExperience Measuring Domain Gap Between Synthetic and Real DataExperience With Generative Image, Video, or Large Language Model APIsFamiliarity With Image and Video File Formats and Metadata StandardsExperience With Batch Compute Platforms or Job Schedulers

Full job description

Weekly Hours: 40

Role Number: 200686470-3577

Summary

Apple's machine learning features are only as good as the data behind them, and our team builds the platform that lets teams create data that fills the gaps real-world collection can't, provides privacy-preserving ways to train and evaluate our models safely, and helps us ensure our products have seen diverse inputs to generalize properly across the many parts of the world where we operate.

In the AI & Machine Learning (AIML) organization, our team builds the self-service tools and platform that turn ideas, text, and images into large-scale generated datasets. We don't produce the datasets ourselves: ML engineers, data scientists, designers, feature teams, ML data operations teams, and Safety and Responsible AI teams across Apple use our platform to generate their own. That data is only valuable if it matches the real world closely enough to train on and to evaluate against, proving our models and features meet Apple's quality bar before they reach customers. As a software engineer on the platform, you'll sit between that technology and the people who depend on it: understanding what they want to create, building the features that make it possible, and helping them get the most from generative AI. It's a hands-on engineering role for someone who enjoys the people side as much as the code.

Description

You'll spend your days moving between engineering and partnership. Synthetic data here does far more than fill gaps: it lets teams build privacy-preserving digital humans, represent locales and domains that real-world collection underserves, and experiment in days instead of waiting on slow, costly data collection. It also gives Safety and Responsible AI red teams the data they need to stress-test models against misuse and edge cases. One day you might be integrating a new image or video generation model and tuning how it runs on GPUs; the next, you might be helping a design team get their first dataset running on the platform, or working with a feature team on model ablations that measure what synthetic data adds to training and evaluation.

You'll help teams understand where synthetic data differs from the real data their models see, whether that's how images look or subtler differences in metadata, distributions, and labels, and then build what closes those gaps. Your core partners are ML engineers and data scientists, but you'll also work with designers, artists, and others who are newer to machine learning, and help make these systems understandable and approachable for them. You'll own features end to end, from shaping the request with partner teams to validating the result on production infrastructure.

Our team builds with AI coding agents every day, and you'll help shape how we use them well. You don't need a research background in generative models; you'll learn the models on the job. What matters most is strong engineering judgment and the ability to bring people along.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or a related field, or 3 years of equivalent work experience.

  • 3+ years of experience building and operating production software in Python.

  • Experience debugging distributed systems or data pipelines in production, using logs, metrics, and task state to find root causes.

  • Experience deploying machine learning models on GPU infrastructure, including dependency management and GPU memory sizing.

  • Experience designing input validation and configuration contracts for systems where late failures are costly.

  • Demonstrated ability to explain machine learning concepts and tradeoffs to non-technical partners such as designers or product managers.

  • Proven, hands-on use of agentic AI coding tools in day-to-day software development, including judging where they are effective and where their output needs human verification.

Preferred Qualifications

  • Experience working directly with partner teams or internal customers to define and ship features.

  • Experience measuring the domain gap between synthetic and real data, including visual statistics and non-visual properties such as metadata and label distributions.

  • Experience with generative image, video, or large language model APIs, including structured output.

  • Familiarity with image and video file formats and metadata standards.

  • Experience with batch compute platforms or job schedulers.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.