Meta logo

Data Center Production Operations Engineer

Meta

Newark, CAJob$111–159K/yrTracked 3w agoSeen in employer's feed 5 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$111–159K/yr
Location
Newark, CA
Work Authorization
Not specified

Job overview

Meta is hiring a Data Center Production Operations Engineer. Meta seeks a Data Center Production Operations Engineer to ensure reliability, efficiency, and scalability of its global data center infrastructure, managing server fleet health, operational processes, and cross‑functional collaborations to support billions of users worldwide.

Key focus areas include Manage and maintain large‑scale server fleets across data center environments, Monitor production systems health using observability tooling and telemetry data, and Develop and refine operational runbooks, escalation procedures, and incident response playbooks.

Important skills include Hardware Triage, Failure Analysis, Coordinating Repair And Replacement Workflows, Monitoring Production Systems Health, Observability Tooling, and Telemetry Data. Preferred (not required): Capacity Planning and Health Checks.

Skills & qualifications

RequiredNice to have

Skills

Hardware TriageFailure AnalysisCoordinating Repair and Replacement WorkflowsMonitoring Production Systems HealthObservability ToolingTelemetry DataDeveloping Operational RunbooksEscalation ProceduresIncident Response PlaybooksCollaborate With Hardware EngineeringNetwork OperationsCapacity PlanningServer DeploymentDecommissioningLifecycle TransitionsAnalyze Failure TrendsOperational Data AnalysisRoot Cause AnalysisCorrective ActionAutomation InitiativesServer ProvisioningHealth ChecksFleet Management WorkflowsAI-Integrated ToolingProcess ImprovementsOperational EfficiencyReduce Mean Time to ResolutionCommunicate Infrastructure StatusIncident TimelinesRisk AssessmentsCommunicationValidating Server Acceptance CriteriaCoordinating With Data Center TechniciansHardware Bring-UpCommissioningIdentifying Gaps in Monitoring CoverageOperational ToolingFleet VisibilityProduction ReliabilityServer Hardware ComponentsCPUsMemoryStorageNetwork Interface CardsHands-on TroubleshootingFailure DiagnosisObservability PlatformsTracking Fleet HealthIdentifying Anomalies

Qualifications

24/7 on-Call RotationTravel Up to 15%6+ Years of Experience in Data Center Operations6+ Years of Experience in Production Infrastructure Engineering6+ Years of Experience With Server Hardware ComponentsExperience Using Systems Monitoring and Observability PlatformsExperience Developing or Improving Operational ProcessesExperience Collaborating With Hardware Engineering, Network, and Capacity TeamsExperience Contributing to Post-Incident ReviewsExperience With Scripting Languages Such as Python or BashFamiliarity With Server Firmware ManagementBackground in Capacity Planning or Hardware Acceptance Testing Processes

Full job description

Summary:

Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. In this role, you will be responsible for the day-to-day operational health of server fleets and production systems that underpin Meta's family of apps and services. You will work at the intersection of hardware lifecycle management, systems reliability, and operational process improvement, ensuring that production environments meet the demands of billions of users worldwide.

Required Skills:

Data Center Production Operations Engineer Responsibilities:

  1. Manage and maintain large-scale server fleets across data center environments, including hardware triage, failure analysis, and coordinating repair and replacement workflows

  2. Monitor production systems health using observability tooling and telemetry data to proactively identify and resolve infrastructure anomalies before they impact service availability

  3. Develop and refine operational runbooks, escalation procedures, and incident response playbooks specific to data center server environments

  4. Collaborate with hardware engineering, network operations, and capacity planning teams to support server deployment, decommissioning, and lifecycle transitions

  5. Analyze failure trends and operational data to identify systemic issues in server hardware or firmware, and drive root cause analysis and corrective action

  6. Contribute to automation initiatives that reduce manual toil in server provisioning, health checks, and fleet management workflows, including leveraging AI-integrated tooling

  7. Partner with cross-functional teams to evaluate and implement process improvements that increase operational efficiency and reduce mean time to resolution for production incidents

  8. Communicate infrastructure status, incident timelines, and risk assessments to engineering and operations stakeholders through clear written and verbal updates

  9. Support capacity readiness activities by validating server acceptance criteria and coordinating with data center technicians during hardware bring-up and commissioning

  10. Identify gaps in monitoring coverage or operational tooling and propose solutions that improve fleet visibility and production reliability

  11. Participate in 24/7 on-call rotation

  12. Ability to travel up to 15% of the time

Minimum Qualifications:

Minimum Qualifications:

  1. 6+ years of experience in data center operations, site operations, or production infrastructure engineering supporting large-scale server environments

  2. 6+ years of experience with server hardware components including CPUs, memory, storage, and network interface cards, including hands-on troubleshooting and failure diagnosis

  3. Experience using systems monitoring and observability platforms to track fleet health, identify anomalies, and drive incident resolution in production data center environments

  4. Experience developing or improving operational processes, runbooks, or automation scripts to support server fleet management at scale

  5. Experience collaborating with hardware engineering, network, and capacity teams to coordinate infrastructure deployments and lifecycle activities

Preferred Qualifications:

Preferred Qualifications:

  1. Experience contributing to post-incident reviews and translating findings into durable operational improvements that reduce recurrence across a server fleet

  2. Experience with scripting languages such as Python or Bash to automate data center operations tasks including health checks, inventory management, or alerting workflows

  3. Familiarity with server firmware management, BIOS configuration, and out-of-band management interfaces such as IPMI or Redfish in hyperscale data center environments

  4. Background in capacity planning or hardware acceptance testing processes within a large-scale cloud or hyperscale data center organization

Public Compensation:

$111,010/year to $158,995/year + bonus + equity + benefits

Industry: Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].

You've read the whole posting — now see how you match it.