DPU Silicon RAS and Debug Architect
Sunnyvale, CAJob$178–250K/yrSeen todaySeen in employer's feed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Sunnyvale, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Meta's Infrastructure Silicon organization designs custom silicon that powers data center infrastructure, including SmartNICs, IPUs, DPUs, AI accelerators, and networking ASICs, and seeks a DPU RAS and Debug Architect to define reliability, availability, and serviceability architecture for DPU ASICs.
Skills & qualifications
Skills
Qualifications
Full job description
Summary:
Meta's Infrastructure Silicon organization designs custom silicon that powers our data center infrastructure — SmartNICs/IPUs/DPUs, AI accelerators, and networking ASICs.We are looking for a Data Processing Unit (DPU) RAS and Debug Architect to define the reliability, availability, and serviceability architecture for our DPU ASICs. You'll set FIT-rate targets from the intended usages and deployment models, define how errors are detected, corrected, and reported, and define the trace, debug, and performance-monitoring architecture that supports both post-silicon debug and software and firmware debugging. You'll decide what belongs in hardware versus firmware versus software, define how the hardware presents itself to firmware and software, and work hand in hand with the firmware, driver, and RTL teams - carrying your designs from early path-finding through implementation and silicon bring-up.
Required Skills:
DPU Silicon RAS and Debug Architect Responsibilities:
-
Own the RAS and Debug architecture for our DPUs - error detection, correction, and reporting, FIT budgeting, and the trace, debug, and telemetry infrastructure - from early path-finding through silicon bring-up
-
Set FIT-rate targets from the intended usages and deployment models, and budget them across the design - memories, logic, on-chip interfaces, and links
-
Define the error detection, correction, reporting, and containment architecture - parity, ECC/SECDED, poisoning and poison propagation, error logging, and how errors are surfaced to firmware, host software, and platform management
-
Define the trace, debug, and performance-monitoring architecture for post-silicon debug and for software and firmware debugging - on-chip trace, event and counter telemetry, crash and state capture, and the JTAG/debug-access model
-
Decide what belongs in hardware versus firmware versus software, and define how the hardware presents itself to them - register and programming models, error and interrupt models, and trace/telemetry interfaces
-
Work closely with design, DV, and PD teams on feature definition and PPA tradeoff
-
refine architecture to meet the design and PD constraints. Define architecture to maximize DV complexity including defining specific features to help ease DV
-
Support Design, DV, PD and DFT teams to resolve interface and integration issues as they come up. Support post-silicon bring-up
Minimum Qualifications:
Minimum Qualifications:
-
8+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
-
Understanding of RAS concepts: FIT-rate estimation and budgeting, failure modes (including silent data corruption), error detection, correction, and containment, and reliability targets for data center deployments. Familiarity with data center reliability, serviceability, and manageability requirements
-
Understanding of error-protection mechanisms - parity, ECC/SECDED, CRC, data poisoning and poison propagation, and lockstep/redundancy techniques - and their area, latency, and power trade-offs
-
Understanding of memory and interface RAS -- LPDDR/DDR RAS (ECC, on-die ECC, error scrubbing, post-package repair) and PCIe RAS (Advanced Error Reporting, ECRC, link error detection and recovery)
-
Understanding of on-chip debug and trace architectures -- JTAG and debug access, on-chip trace, breakpoints and watchpoints, and crash and state capture (e.g., ARM CoreSight or comparable)
-
Understanding of performance-monitoring architectures - hardware performance counters and event telemetry - and their use in post-silicon and software/firmware debug
-
Familiarity with DFT concepts (scan, MBIST/LBIST, boundary scan) and how they interact with RAS and debug
-
Familiarity with processor ISA debug mechanisms (ARM and/or x86) and instruction/execution trace mechanisms such as ETM
-
Knowledge of relevant industry standards and specifications, such as OCP (server, RAS, and telemetry/manageability specifications), JEDEC (LPDDR/DDR), and PCIe
-
Demonstrated ability to drive analysis independently and influence architectural direction through data
Preferred Qualifications:
Preferred Qualifications:
-
Experience defining FIT budgets and reliability targets for hyperscale data center silicon, including soft-error rate (SER) analysis and mitigation
-
15+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
-
Experience with hardware description languages (e.g., SystemVerilog, VHDL) and simulation environments used in ASIC development flows
-
Experience developing Python-based (or other scripting) automation pipelines for debug, telemetry collection, and data analysis
-
Familiarity with functional-safety standards (e.g., ISO 26262) and reliability qualification methods
-
Experience designing error-reporting and machine-check/AER architectures and the firmware error-handling flows built on them
-
Experience designing on-chip trace and debug subsystems (e.g., ARM CoreSight) and the associated post-silicon debug tooling
Public Compensation:
$178,000/year to $250,000/year + bonus + equity + benefits
Industry: Internet
Equal Opportunity:
Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior DFT Engineer - Post SiliconNVIDIA · Santa Clara, CA (Hybrid) · $168–311K/yrPosted 2 days agoPosted 2 days ago
Senior Silicon Circuit Co-Design EngineerNVIDIA · Santa Clara, CA · $168–265K/yrPosted 2w agoPosted 2w ago
Silicon SoC ArchitectIntel · Santa Clara, CA (Hybrid) · $221–361K/yrPosted 3w agoPosted 3w ago
Principal Hardware Architect – ASIC & System Integration, Cisco Silicon One (Hybrid)Cisco · San Jose, CA (Hybrid) · $234–330K/yrPosted 1w agoPosted 1w ago
Senior Manager, Silicon Speed Productization — Silicon Co-Design GroupNVIDIA · Santa Clara, CA (Hybrid) · $232–368K/yrPosted 4 days agoPosted 4 days ago
You've read the whole posting — now see how you match it.