Meta logo

Cloud Production Systems Engineer

Meta

Los Lunas, NMJob$173–245K/yrSeen 2 days agoSeen in employer's feed today

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$173–245K/yr
Location
Los Lunas, NM
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Meta seeks an experienced Cloud Production Systems Engineer to join the Data Center Operations team, focusing on cloud automation, data analysis, and tooling to improve server fleet uptime and efficiency across global data centers.

Skills & qualifications

RequiredNice to have

Skills

AWSGCPCoreWeaveNeoCloudPythonBashPHPCC++LinuxWeb ServersLoad BalancersRelational DatabasesStorage SystemsMessaging SystemsData AnalysisVisualization ToolsConfiguration ManagementInfrastructure‑as‑CodeDistributed Systems MonitoringAlertingAutomated Remediation Pipelines

Qualifications

Bachelor's Degree in Computer Science or Computer Engineering or Relevant Technical Field or Equivalent Practical Experience7+ Years Production Systems Engineering or Infrastructure Engineering or System Software Development for Large-Scale Cloud Environments7+ Years Hardware Lifecycle Management or Fleet Automation or Cloud Data Center Operations Systems Spanning Compute Storage or Networking InfrastructureExperience Developing Systems Software or Automation Tooling in Python Bash PHP C or C++ for Linux-Based Production Environments at ScaleExperience With Configuration and Maintenance of Production Systems Including Web Servers Load Balancers Relational Databases Storage Systems and Messaging SystemsExperience Communicating Technical Designs and Cloud Infrastructure Decisions Through Written Documentation and Cross-Functional Stakeholder AlignmentExperience Designing or Operating Cloud Configuration Management and Infrastructure-as-Code Systems for Large Heterogeneous Hardware FleetsExperience With Data Analysis and Visualization Tools Used to Prioritize Fleet Health Initiatives and Drive Operational Decision-MakingExperience Supporting Global Multi-Site Data Center Infrastructure Deployments Including Hardware Qualification and Regional Rollout CoordinationFamiliarity With Distributed Systems Monitoring Alerting and Automated Remediation Pipelines at Hyperscale

Full job description

Summary:

Meta is seeking an experienced Cloud Production Systems Engineer to join the Data Center Operations team. Our data centers and the cloud infrastructure powering tens of thousands of servers forms the foundation upon which Meta's rapidly scaling services operate. Meta is at the leading edge of the global data center industry in both design and cloud operations. This role requires a forward-thinking systems professional with deep experience leveraging diverse software tools to identify cloud automation solutions for complex operational challenges with our Cloud partnerships (AWS, GCP, CoreWeave, and other NeoCloud). The ideal candidate performs deep data analysis to prioritize server repair automation in a hyperscale cloud environment, drives solutions through code, and collaborates effectively with globally distributed teams through clear written communication.

Required Skills:

Cloud Production Systems Engineer Responsibilities:

  1. Identify and root cause systemic issues across the cloud server fleet and drive resolutions to maximize uptime and utilization by leveraging hardware failure data and diagnostic telemetry

  2. Write, review, and maintain code for diagnostic and cloud automation tooling that supports quality and efficient delivery of production servers at hyperscale

  3. Own and develop diagnostic tooling requirements that enable frontline operations teams to efficiently manage and repair the cloud server fleet

  4. Drive the escalation process for Data Center Operations to identify, root cause, and resolve complex cloud tooling and hardware issues affecting fleet health

  5. Execute operational validation and verification activities for new cloud product integration into the production environment

  6. Collaborate with cross-functional cloud tooling teams to provide an operations-centric perspective on open issues and contribute to their development roadmaps

  7. Perform deep data analysis to prioritize cloud automation opportunities for server repair workflows in a large-scale, heterogeneous hardware environment

  8. Build cross-functional relationships and influence policies and procedures to improve global cloud data center operations consistency and efficiency

  9. Mentor other engineers on evaluating and resolving cloud fleet issues and defining improvements to tools and operational processes

  10. Travel up to 25% to support global cloud data center operations and new site deployments

Minimum Qualifications:

Minimum Qualifications:

  1. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

  2. 7+ years of experience in production systems engineering, infrastructure engineering, or system software development for large-scale cloud environments

  3. 7+ years of experience with hardware lifecycle management, fleet automation, or cloud data center operations systems spanning compute, storage, or networking infrastructure

  4. Experience developing systems software or automation tooling in Python, Bash, PHP, C, or C++ for Linux-based production environments at scale

  5. Experience with the configuration and maintenance of production systems, including web servers, load balancers, relational databases, storage systems, and messaging systems

  6. Experience communicating technical designs and cloud infrastructure decisions through written documentation and cross-functional stakeholder alignment across engineering and operations teams

Preferred Qualifications:

Preferred Qualifications:

  1. Experience designing or operating cloud configuration management and infrastructure-as-code systems for large heterogeneous hardware fleets

  2. Experience with data analysis and visualization tools used to prioritize fleet health initiatives and drive operational decision-making

  3. Experience supporting global, multi-site data center infrastructure deployments, including hardware qualification and regional rollout coordination

  4. Familiarity with distributed systems monitoring, alerting, and automated remediation pipelines at hyperscale

Public Compensation:

$173,000/year to $245,000/year + bonus + equity + benefits

Industry: Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.