Zaimler logo

ML Infrastructure Engineer

Zaimler

San Mateo, CAFull-timeNo compensation foundPosted 5mo agoVerified open 6 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
San Mateo, CA
Schedule
Full-time
Work Authorization
Visa required • Visa sponsorship

Requirements

Credentials this posting asks for.

Doctorate

Job overview

Zaimler is hiring a ML Infrastructure Engineer. Zaimler is building a context infrastructure platform that enables autonomous AI agents to reason over fragmented enterprise data by automatically discovering domain knowledge and mapping relationships, delivering real‑time inference through knowledge graphs for precision‑scale operations.

Key focus areas include Set up and scale inference/Ray Serve for ML and LLM model serving, integrated with data analysis and agent workflows, Scale agent GPU infrastructure for concurrency and efficiency across multiple agent workloads, and Optimize and improve the engine builder and model server that power scalable agent orchestration.

Successful candidates bring PhD In CS, ML, Or Related Field and MS With 4+ Years Relevant Industry Experience. Important skills include LLM Optimization, Quantization, Memory Layout, Serving Performance, LLM Source Code Debugging, and Runtime Libraries Debugging. Preferred (not required): Feature Store Design, Feature Store Management, GPU Cluster Management, and GPU Cluster Optimization.

Skills & qualifications

RequiredNice to have

Skills

LLM OptimizationQuantizationMemory LayoutServing PerformanceLLM Source Code DebuggingRuntime Libraries DebuggingRustC++PythonAlgorithmic FundamentalsData Structures and AlgorithmsComplexityDistributed SystemsModel Serving InfrastructurevLLMBasetenTritonML Pipelines Setup and ScalingFeature Store DesignFeature Store ManagementGPU Cluster ManagementGPU Cluster OptimizationOpen-Source ML Infrastructure ContributionsLLM Tooling ContributionsRayONNXTensorRTTransformer Internals UnderstandingAttention Mechanisms UnderstandingKV CacheRay ServeAIBrixGPU InfrastructureModel ServingInference StackBuild Scalable ML AI Platforms

Qualifications

3+ Years Relevant Experience

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match

Full job description

About zaimler AI agents can't reason over data they don't understand. Enterprise data today is fragmented across dozens of systems with no shared context, meaning, or structure, and that's why most enterprise AI is failing. The shift from copilots to autonomous agents is creating an entirely new infrastructure layer, and we're building it. zaimler is the context infrastructure for the agentic era: a platform that automatically discovers domain knowledge, maps relationships, and gives AI agents the semantic understanding to operate with precision at scale. Imagine knowledge graphs that support real-time inference, built for systems that need to reason, not just retrieve. zaimler was founded by Biswajit Das (ex-VP Engineering, Truera), a Data Infra veteran and former Chief Architect at Visa, and Sofus Macskassy (ex-Director of Engineering, LinkedIn), who built one of the largest knowledge graphs in production in the industry at LinkedIn. We're growing and deploying with major enterprises across insurance, travel, and technology. If you want to build infrastructure that the next decade of enterprise AI runs on, we'd love to talk. About the Role You'll own our inference and model-serving infrastructure end to end. This isn't a research role. It's a build role: you're setting up and scaling the systems that let our agents actually run in production, fast and reliably, at increasing concurrency.

You report to Sofus and work closely with our ML and infra teams.

What You’ll Own

  • Set up and scale inference/Ray Serve for ML and LLM model serving, integrated with our data analysis and agent workflows
  • Scale agent GPU infrastructure for concurrency and efficiency across multiple agent workloads
  • Optimize and improve the engine builder and model server that power scalable agent orchestration What You Need
  • Proven ability to build scalable ML/AI platforms from scratch, end-to-end, for production use cases. You've owned a zero-to-one build before, or can show you're capable of it
  • Deep understanding of the inference stack: vLLM, KV cache, and the optimization layers underneath model serving
  • Experience building distributed systems for AI/ML workloads at scale, connecting them to real product or vertical integrations
  • 3+ years of relevant experience. We care about capability, not tenure Nice to Have
  • Ray / Ray Serve experience
  • Familiarity with AIBrix Why Join
  • A rare chance to shape both company and product direction as an early team engineer
  • Work alongside engineers and researchers from LinkedIn, Visa, Meta, and Branch
  • Onsite culture in San Mateo, built for deep collaboration and high-velocity building
  • Full benefits (medical, dental, vision, 401k)
  • We sponsor H-1B visas and assist with immigration We value builders over résumés. If this role excites you but you don't check every box, we still want to hear from you. zaimler is an equal opportunity employer.

You've read the whole posting — now see how you match it.

ML Infrastructure Engineer at Zaimler | Olive Jobs