System Engineer, AI Accelerator Module Management
Menlo Park, CAJob$173–245K/yrSeen todaySeen in employer's feed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Menlo Park, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
The System Engineer will design, implement, and support hardware management for Meta’s custom AI accelerators. The role focuses on specifying ASIC management and debug functionality, integrating device management with fleet systems, and developing tools and concepts to improve accelerator reliability and visibility. The engineer will work across silicon, hardware, software, and data-center teams, helping architect solutions and deliver hardware management at scale.
Skills & qualifications
Skills
Qualifications
Full job description
Summary:
The Accelerator Reference Design Team is looking for a System Engineer to design, implement, and support hardware management for custom AI hardware. The ARD team is working on the design and implementation of hardware modules for Meta's custom AI accelerators.Meta is developing large-scale AI/HPC clusters using custom-designed AI accelerators. In this role, you will have a unique opportunity to shape the future AI/HPC of Meta by specifying technical requirements for device management, driving specifications and designs for integrating device management into fleet management, and steering the industry and ecosystem partner direction.The ideal candidate will work in a cross-functional engineering environment, prioritizing competing workstreams based on impact, deadlines, and stakeholder needs. They will have experience in solving complex device management problems in high-performance computing that span across silicon, hardware, and software. They will also have direct experience in hardware design and management of complex ASICs. A successful candidate will have experience working on hardware management across rack, server, and ASIC levels. The candidate will have experience with current industry practices and standards for device management and debug. The position requires a lead developer to architect solutions, debug complex issues, define scope for open-ended technical challenges, and deliver hardware management at scale. The candidate will work closely with cross-functional stakeholders and partners who are on the front-line of developing Meta's custom ASICs for AI.The Accelerator Reference Design Team designs, builds, brings-up, tests and integrates hardware systems that power Meta's custom AI silicon platforms, deployed in data centers worldwide. In this role, you will help design and build open and efficient AI platforms deployed at scale.
Required Skills:
System Engineer, AI Accelerator Module Management Responsibilities:
-
Work as part of the Accelerator Reference Design Team to specify, develop, and integrate device management for Meta's custom AI hardware
-
Collect requirements and develop specifications for AI/HPC ASIC management and debug functionality
-
Develop and maintain code, to collect, analyze, and interpret data for accelerator fleet statistics
-
Collaborate with cross-functional teams to develop detailed management specifications for our silicon and platforms, focusing on AI accelerator ASIC management
-
Develop concepts and proof-of-concept experiments to improve reliability and visibility of new and existing AI accelerators
-
Specify and Design the module and ASIC management of Meta's custom AI silicon so that it integrates well with the data centers and the fleet
Minimum Qualifications:
Minimum Qualifications:
-
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
-
8+ years of hands-on experience in implementing management solutions for computer server systems
-
4+ years of designing or specifying custom silicon management
-
Proficiency in C++ and/or Python, or equivalent
-
Experience with general understanding of architectural trade-offs in hardware and software across cost, performance, and reliability
-
Experience working effectively as an individual and in a multidisciplinary team
-
Troubleshooting experience across software, firmware, hardware, and network problems
Preferred Qualifications:
Preferred Qualifications:
-
6+ years of experience in device management for high-performance custom silicon
-
Core domain knowledge in servers and networking with experience in debugging operational problems
-
Datacenter level tooling experience for management
-
Experience in developing both custom hardware and software for AI/HPC
-
Experience with system performance analysis, debug, and optimization practices
-
10+ years of hands-on experience in implementing device management for computer server systems
Public Compensation:
$173,000/year to $245,000/year + bonus + equity + benefits
Industry: Internet
Equal Opportunity:
Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior System Software Engineer - GPU Power and Performance ManagementNVIDIA · Santa Clara, CA · $184–288K/yrPosted 2w agoPosted 2w agoSenior Staff AI Accelerator Performance ArchitectCerebras · Sunnyvale, CA · $175–275K/yrPosted 2w agoPosted 2w ago
Senior ML Accelerator Engineer - GPUGeneral Motors · Sunnyvale, CA (Hybrid) · $170–258K/yrPosted 1w agoPosted 1w ago
Senior/Staff System Research Engineer – LLM Inference OptimizationSnowflake Inc. · Bellevue, WA (Hybrid) · $236–330K/yrPosted 2w agoPosted 2w ago
Senior System Software Engineer - Embedded ControllerNVIDIA · Santa Clara, CA · $152–242K/yrPosted 2w agoPosted 2w ago
You've read the whole posting — now see how you match it.