Fractile logo

ML Runtime Engineer

Fractile

Bristol, United KingdomJobNo compensation foundPosted 4mo agoVerified open 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Bristol, United Kingdom
Work Authorization
Not specified

Job overview

Fractile’s ML Runtime Engineer will join the Developer Experience team to integrate AI accelerators with leading inference frameworks, research and build KV‑cache management solutions, and design a scalable reference inference engine focused on transformer architectures, while contributing to a cutting‑edge software stack for AI inference systems.

Skills & qualifications

RequiredNice to have

Skills

Inference EngineSoftware EngineeringML InferenceMulti‑User ServingPaged AttentionvLLMSGLangRust

Qualifications

Degree in Computer Science or Related Field

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match
Paid Time Off

Full job description

ML Runtime Engineer Location: Bristol / London About Fractile Fractile was founded in 2022 on the bet that, eventually, the world’s most capable AI systems would be limited in their impact by the time taken to produce useful outputs. We bet everything on the logical conclusion: that the only way to truly unlock this latent value, to make speed viable at scale, was to radically re-invent the hardware that we run our frontier AI models on. Ever since, we have been building chips and systems that tackle this problem: how to efficiently generate output at thousands of tokens per second, while handling the complexity and capacity challenges of operating large models at very long contexts. The workloads that push to the limits of the current frontier are already transformational; it is the technical and economic limits on inference speed that are constraining progress. The defining work of the 21st century will be marked by the engine of inference delivering immense and diffuse chains of intellectual inquiry, in drug discovery, in software engineering, in materials discovery, in any field where progress is driven by deep reasoning and intelligence to resolve complex problems. About the Software organisation at Fractile Developer Experience sits within the Software organisation at Fractile, which is responsible for developing a full software stack for our groundbreaking AI inference systems. That's everything from ML compilers, device drivers and systems firmware, application level runtime and ecosystem integrations, ML and compute libraries, great developer tooling and a full portfolio of simulators, through to datacenter scale workload deployment solutions. At Fractile, we know that a fantastic software stack is a critical and central part of any AI inference solution and it sits at the heart of everything we're doing. About the team & role The Role About the Software organisation at Fractile Developer Experience sits within the Software organisation at Fractile, which is responsible for developing a full software stack for our groundbreaking AI inference systems. That's everything from ML compilers, device drivers and systems firmware, application level runtime and ecosystem integrations, ML and compute libraries, great developer tooling and a full portfolio of simulators, through to datacenter scale workload deployment solutions. At Fractile, we know that a fantastic software stack is a critical and central part of any AI inference solution and it sits at the heart of everything we're doing. About the team and role The ML Runtime team is responsible for integrating Fractile's AI accelerators with the latest inference frameworks and building the runtime stack that makes them fly. We work on genuinely hard problems — KV cache management, scalable multi-user inference, and the internals of transformer model execution — alongside a collaborative team that values curiosity and rigour equally. As an ML Runtime Engineer you will integrate Fractile's AI acceleration hardware with leading inference engines including vLLM and SGLang, research and build proof-of-concept KV cache management implementations tailored to our hardware, and work closely with the broader runtime team to design and build a scalable reference inference engine. You will focus primarily on the transformer ML architecture and share your expertise to help shape the direction of our runtime stack. About you You have solid experience with ML inference at scale, including multi-user serving, and a deep understanding of paged attention and inference engines such as vLLM. You are familiar with the key components of the ML software ecosystem and bring strong software engineering skills with an instinct for clean, maintainable systems. You care about depth of knowledge and have a genuine interest in the problem space — not just in shipping, but in understanding why things work the way they do. Key Requirements

  • Solid experience with ML inference at scale, including multi-user serving
  • Deep understanding of paged attention and inference engines such as vLLM
  • Familiarity with key components of the ML software ecosystem
  • Strong software engineering skills and an instinct for clean, maintainable systems

Nice to Have

  • Experience with Rust
  • Having built your own inference engine from scratch
  • A degree in Computer Science or a related field

What We Offer

  • Competitive salary: A competitive salary reflective of your experience and the specialist nature of the role.
  • Equity & Ownership: meaningful equity so everyone shares in the value creation
  • Benefits: Private Medical, Dental and Vision, Contributory Pension, 25 Days holiday plus bank holidays and Life/Critical Illness Insurance.
  • Diverse & fun office: we believe the hardest problems get solved by the broadest range of minds. We are committed to Equal Employment Opportunity through attracting and retaining a diverse team and building an inclusive environment.

Fractile is seeking to increase the clock speed of global progress, one chip at a time. We’ve recently raised $220M from investors including Founders Fund and Accel and our most important work lies ahead. Join us! Export controls Our work involves technologies subject to UK, US and other international export control regulations. Certain roles may require additional eligibility checks to ensure compliance with applicable law. We'll be transparent about this throughout the hiring process.

You've read the whole posting — now see how you match it.