Apple logo

Sr Machine Learning Engineer, Proactive

Apple

Santa Clara, CAFull-timeSeen 1 day agoSeen in employer's feed 1 day ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Santa Clara, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Requirements

Credentials this posting asks for.

Master's degree

Job overview

Apple is seeking a Senior Machine Learning Engineer to design, train, fine‑tune, optimize and deploy large language models, semantic retrieval systems and ranking models that power personalized, context‑aware experiences across its ecosystem while preserving privacy.

Skills & qualifications

RequiredNice to have

Skills

PythonC/C++PytorchJaxTensorflowTransformersEmbeddingsRepresentation LearningNeural RankingLarge Language ModelsFoundation ModelsKnowledge DistillationModel CompressionQuantizationPruningOn‑Device Machine LearningEdge AiMobile InferenceRetrieval‑Augmented GenerationVector SearchEmbedding RetrievalNeural RerankingSemantic SearchQuery UnderstandingQuery RewritingIntent ClassificationPersonalized RetrievalLearning‑to‑RankRecommendation ModelsBertT5LlamaGemmaMistralMultimodal Foundation ModelsAgentic AiAgentic RetrievalAi Quality MetricsEvaluation PipelinesLarge‑Scale Production Search

Qualifications

Master Degree in Computer Science, Machine Learning, Artificial Intelligence or Related Field5+ Years of Industry or Research Experience Developing Machine Learning SystemsBackground in Machine Learning, Deep Learning, Natural Language Processing, Information Retrieval, Search, Recommender Systems, or Generative AiExperience Training, Fine‑Tuning, or Deploying Transformer‑Based Models and Large Language ModelsExperience With Modern Deep Learning Architectures and Techniques Including Transformers, Embeddings, Representation Learning, and Neural RankingProgramming Skills in Python and/or C/C++ With Experience Building Production‑Quality Software Using Modern Machine Learning Frameworks Such as Pytorch, Jax, or TensorflowAbility to Work Onsite in Cupertino, California, in Accordance With Apple’s Applicable Work PoliciesMaster’s or Ph.D. In Computer Science, Machine Learning, Artificial Intelligence, or Related FieldExperience Optimizing Machine Learning Models for Resource‑Constrained Environments Including Knowledge Distillation, Model Compression, Quantization, and PruningExperience With on‑Device Machine Learning or Edge Ai, or Mobile Inference Frameworks Including Optimizing Models for Latency, Memory, Compute, and Power ConstraintsExperience Distilling Capabilities From Large Foundation Models Into Small Language Models or Task‑Specific Models for Efficient InferenceExperience Building Retrieval‑Augmented Generation, Vector Search, Embedding Retrieval, Neural Reranking, or Semantic Search Systems

Full job description

Weekly Hours: 40

Role Number: 200672628-3760

Summary

At Apple, machine learning powers experiences that anticipate what people need before they ask. We're looking for a Senior Machine Learning Engineer to help build the next generation of intelligent search and AI experiences technology that understands user intent, context, and personal information while preserving privacy. In this role, you'll design, train, fine-tune, optimize, and deploy large language models, semantic retrieval systems, and ranking models that power relevant, personalized, and context-aware experiences across Apple's ecosystem.

Description

You'll design, train, fine-tune, and optimize transformer-based language models and foundation models for efficient on-device deployment, and build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems that improve search quality and AI-powered experiences. You'll develop models for query understanding, intent prediction, personalization, retrieval, and ranking, while researching new approaches to LLM fine-tuning, knowledge distillation, model compression, quantization, and low-latency inference. You'll explore techniques for adapting large foundation models into smaller, highly capable models that can operate efficiently under on-device memory, compute, power, and latency constraints. You'll partner with engineers, researchers, product managers, and designers to bring new AI capabilities from research into production, driving technical strategy and leading projects from early exploration through large-scale deployment. This is an opportunity to explore new applications of foundation models, multimodal AI, agentic retrieval, and personalized intelligence, shaping the next generation of proactive and intelligent user experiences.

Minimum Qualifications

  • Master degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.

  • 5+ years of industry or research experience developing machine learning systems.

  • Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.

  • Experience training, fine-tuning, or deploying transformer-based models and large language models.

  • Experience with modern deep learning architectures and techniques, including transformers, embeddings, representation learning, and neural ranking.

  • Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow.

  • Ability to work onsite in Cupertino, California, in accordance with Apple's applicable work policies.

Preferred Qualifications

  • Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.

  • Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.

  • Experience with on-device machine learning or edge AI, or mobile inference frameworks, including optimizing models for latency, memory, compute, and power constraints.

  • Experience distilling capabilities from large foundation models into small language models or task-specific models for efficient inference.

  • Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.

  • Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.

  • Experience working with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.

  • Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.

  • Experience building large-scale production search, recommendation, personalization, or generative AI systems.

  • Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.

  • Strong understanding of the tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.

  • Ability to prototype new ideas, conduct rigorous experiments, solve ambiguous technical problems, and translate research advances into production-quality machine learning solutions.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.