
Software Engineer, Cloud
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
Ollama is hiring a Software Engineer, Cloud. Ollama is seeking a Software Engineer, Cloud to build and scale its cloud inference platform. This role involves working on high-throughput, low-latency distributed systems, including inference serving, GPU fleet management, routing, and metering. The engineer will contribute to the platform that supports Pro, Max, Team, and Enterprise customers, processing trillions of tokens.
Key focus areas include Build and scale the inference platform that serves every request from ollama.com, Design the routing and capacity layer that places workloads across GPUs and regions, and Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and tiering.
Successful candidates bring Deep Experience With High‑Throughput Low‑Latency Distributed Systems, Comfortable With Cost/Performance Tradeoffs At Scale, and Experience With Kubernetes GPU Scheduling Or Inference Infrastructure. Important skills include High-Throughput Distributed Systems, Low-Latency Distributed Systems, Inference Serving, Traffic Routing, Real-Time Data Pipelines, and Large-Scale APIs. Preferred (not required): Building Inference Platform, GPU Fleet Management, and Billing/Metering For AI Service.
Skills & qualifications
Skills
Qualifications
Full job description
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
ABOUT THE ROLE
You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.
WHAT YOU'LL DO
-
Build and scale the inference platform that serves every request from ollama.com http://ollama.com.
-
Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.
-
Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.
-
Build the reliability, observability, and cost controls for our team and customers
YOU MAY BE A FIT IF
-
You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.
-
You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.
-
You've worked with Kubernetes, GPU scheduling, or inference infrastructure.
-
You think in terms of reliability, SLOs, and honest capacity planning.
-
Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.
You've read the whole posting — now see how you match it.