Meta logo

Production Systems Engineer, Sustaining

Meta

Menlo Park, CAJob$144–204K/yrSeen 1 day agoSeen in employer's feed 1 day ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$144–204K/yr
Location
Menlo Park, CA
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Meta seeks an experienced Production Systems Engineer for its Release to Production team, which supports the hardware lifecycle of servers from pre-production testing through production monitoring and remediation. The engineer will support AI and HPC infrastructure, develop and execute system test suites, diagnose hardware and software issues, and coordinate with internal teams and external vendors to improve sustaining practices and test quality.

Skills & qualifications

RequiredNice to have

Skills

Production Hardware SupportCPU Systems DeploymentHardware and Software Co-DesignObject-Oriented ProgrammingPythonC/C++Server and Network Datacenter SystemsCPU Systems Co-DesignLarge-Scale Systems SupportAI/HPC HardwareHardware ConfigurationGPUMemoryNetwork ConfigurationVM EnvironmentsPerformance OptimizationNCCLPyTorchCUDA

Qualifications

Bachelor's in Computer Science, Computer Engineering, or Related Field or Equivalent Practical Experience6+ Years Hardware Systems Experience8+ Years CPU Systems Experience8+ Years Production Support Experience

Full job description

Summary:

Meta is seeking an experienced Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services. The RTP team is responsible for the Hardware Lifecycle of all Meta servers including pre-production hands-on system and hardware debugging and stress testing, enabling production-ready system monitoring, automated provisioning and automated remediation of issues. RTP Engineers work closely with hardware designers, system manufacturers, component vendors, capacity engineering, production engineering, Meta services, and data center operations teams to test systems before release to our production data centers, and to track the health and life cycle of servers in production.

Required Skills:

Production Systems Engineer, Sustaining Responsibilities:

  1. Develop robust, scalable, reliable practices for supporting AI and HPC infrastructure at scale

  2. Interface with external vendors and internal hardware, mechanical, power, thermal, manufacturing and software engineers to understand system architecture to develop and execute the test suites for various architectures

  3. Proactively create experiments and tooling to detect and diagnose hardware/firmware/software health issues

  4. Implement sustaining workflows across hardware and software stacks, develop and communicate sustaining practices internally

  5. Troubleshoot, diagnose and root cause of system failures and isolate components and failure scenarios while working with internal and external stakeholders

  6. Drive necessary discussions with external and internal teams on test specification and methodologies to improve test quality on an ongoing basis

Minimum Qualifications:

Minimum Qualifications:

  1. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

  2. 6+ years of experience in hardware systems technologies or supporting production hardware at scale

  3. Experience in deploying and productionizing CPU systems and/or related components at scale

  4. Experience in software and hardware co-design for hyperscale systems

  5. Experience in object oriented programming (e.g., Python, C/C++)

  6. Experience in different server and network datacenter systems

Preferred Qualifications:

Preferred Qualifications:

  1. 8+ years of experience in working with CPU systems, including hardware and software components, co-design

  2. 8+ years of experience in providing production support for large-scale systems

  3. Knowledge of AI/HPC hardware requirements and specifications (e.g., configuring hardware components, GPU, memory, network for AI/HPC workloads)

  4. Experience in developing or debugging CPU systems in a VM environment, performance optimizations, including familiarity with relevant tools, libraries, and frameworks (e.g., NCCL, PyTorch, CUDA)

Public Compensation:

$144,000/year to $204,000/year + bonus + equity + benefits

Industry: Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.