Meta logo

Software Engineer, GenAI Frameworks - MTIA

Meta

Menlo Park, CAJob$184–257K/yrSeen 1 day agoSeen in employer's feed 1 day ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$184–257K/yr
Location
Menlo Park, CA
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Meta is hiring a Software Engineer to own a key area of generative AI inference on its MTIA accelerator platform. The role spans inference frameworks, distributed inference, graph-mode execution and compilation, and PyTorch integration. The engineer will set technical direction, lead multi-quarter programs across teams, and own accuracy, stability, performance, and test coverage for production workloads on custom silicon.

Skills & qualifications

RequiredNice to have

Skills

Generative AI InferenceModel ServingPythonC++RustLow-Level Systems CodePerformance-Critical ProgrammingPrefill OptimizationDecode OptimizationKV-Cache ManagementKV-Cache CompressionBatchingSchedulingLatency and Throughput TradeoffsGraph-Mode ExecutionHost-Side Latency ReductionMemory Allocation and PlacementTechnical Design LeadershipEnd-to-End Project DeliveryTest StrategyContinuous IntegrationNumerical Accuracy DebuggingAccuracy and Stability ValidationContext and Sequence ParallelismRing AttentionStreaming AttentionHierarchical KV CacheOffloaded KV CacheAI Accelerator PlatformsMachine Learning Framework Backend IntegrationLow-Precision NumericsQuantizationCalibrationAccuracy Error AnalysisDistributed Systems DebuggingAI Workflow OptimizationResponsible AI PracticesRisk AssessmentBias MitigationDistributed Inference

Qualifications

5+ Years GenAI Frameworks or Model Serving

Full job description

Summary:

Meta designs and deploys its own AI systems. MTIA, the Meta Training and Inference Accelerator, is Meta's family of in-house AI accelerator ASICs, deployed in production across Meta's data centers and expanding workload coverage from ranking and recommendation to generative AI. The MTIA Software team is part of the AI & Compute Foundation organization. Because the hardware is ours, the software is ours too: we build the entire stack a chip vendor would normally supply: compiler and toolchain, runtime, kernel libraries, developer tooling, and deep PyTorch integration, co-designed with the silicon teams generation over generation.The GenAI Frameworks org owns the layer where a model meets the machine. We take frontier models, including large language, multimodal and diffusion models, and make them run efficiently on MTIA. Our work determines how fast those models respond, how much context they can afford to hold, how many of them a given amount of silicon can serve, and how quickly a newly released model can be running and trusted in production.We are hiring a Software Engineer to own a key area of GenAI inference on MTIA, spanning serving frameworks, distributed inference, graph-mode execution and compilation, and core PyTorch integration. This is a domain-expert role: you will set technical direction for your area, drive ambiguous multi-quarter programs across team boundaries, and own the quality bar for accuracy, stability, performance, and test coverage on custom silicon.

Required Skills:

Software Engineer, GenAI Frameworks - MTIA Responsibilities:

  1. Serve as technical owner and domain expert for a key GenAI inference framework area: serving runtime, distributed inference, graph-mode execution and compilation, or core PyTorch integration

  2. Lead ambiguous, multi-quarter technical programs end to end: technical design, execution, test strategy and CI, rollout, and production hardening across teams and org boundaries

  3. Design, implement, and ship features in generative AI inference frameworks and the MTIA integration layers beneath them, from prototype through production deployment

  4. Own performance for production inference workloads end to end: profile across frameworks, runtime, compiler, and kernel boundaries, identify where time actually goes, and drive the fix to the correct layer rather than the convenient one

  5. Build and extend the distributed inference substrate: hierarchical KV caching, cross-host transfer paths, disaggregated prefetch and decode, and parallelism strategies for long-context and mixture-of-experts serving

  6. Optimize the runtime and execution path: graph capture and replay, host-side latency, memory allocation and placement, and throughput under real traffic conditions

  7. Enable frontier models on MTIA and validate accuracy against GPU baselines, closing correctness gaps and distinguishing real regressions from stale references

  8. Own the quality bar for your area, ensuring accuracy, stability, latency, and throughput, and test coverage with regression detection that keeps a win from quietly eroding

  9. Partner with kernel, compiler, and silicon teams on hardware/software co-design, producing framework-level evidence that shapes the future silicon while the design can still change

  10. Provide technical leadership: mentor other engineers, raise the design quality through reviews, and communicate decisions clearly through design documents and cross-team reviews

Minimum Qualifications:

Minimum Qualifications:

  1. 5+ years of hands-on experience with generative AI inference or training frameworks, or production model-serving systems

  2. Proficiency in Python, C++, or Rust, including low-level systems code and performance-critical paths

  3. Experience with GenAI inference optimization: prefill and decode optimization, KV-cache management and compression, batching and scheduling, and reasoning about latency and throughput tradeoffs

  4. Experience with runtime-level optimization: graph-mode execution, host-side latency reduction, memory allocation and placement

  5. Demonstrated experience leading technical design and end-to-end delivery of frameworks or infrastructure projects, including test strategies and CI, across team boundaries

  6. Experience debugging numerical accuracy issues and defending accuracy and stability bars, not just performance

Preferred Qualifications:

Preferred Qualifications:

  1. Experience with long-context inference techniques: context and sequence parallelism, ring or streaming attention, hierarchical or offloaded KV cache

  2. Experience with AI accelerator platforms (GPU, TPU, or custom ASICs) and bringing up a new hardware backend in a major machine learning framework

  3. Experience with low-precision numerics and quantization (FP8, block-scaled formats, INT8/INT4), including calibration and accuracy error analysis, debugging tools for distributed systems

  4. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)

  5. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)

  6. Experience with distributed inference or training at scale: tensor, pipeline, and expert parallelism, collective communication primitives, RDMA transports, and multi-host serving, scale-up and scale-out design

  7. Experience with performance optimization in production environments, with a track record of using profiling and tracing to understand and resolve performance bottlenecks

  8. Experience with LLM inference serving stacks, including continuous batching, paged or chunked attention, speculative decoding, prefix caching, disaggregated prefill and decode, or mixture-of-experts routing and load balancing

  9. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies

Public Compensation:

$183,997/year to $257,000/year + bonus + equity + benefits

Industry: Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at [email protected].

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.