Latent logo

Site Reliability Engineer

Latent

San Francisco, CA · HybridFull-time$200–275K/yrPosted 9mo agoStill listed 2w ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$200–275K/yr
Location
San Francisco, CAHybrid
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Requirements

Credentials this posting asks for.

Master's degree

Job overview

Latent is hiring a Site Reliability Engineer. The infrastructure expert enables rapid product development and guarantees 99.9%+ stability and performance of a clinical AI platform, linking operational excellence directly to patient access to life‑saving treatment, while thriving in a high‑energy, in‑office culture.

Key focus areas include Design, implement, and maintain the production environment., Own containerized infrastructure, managing deployment, scaling, and health with Kubernetes and Helm., and Optimize TypeScript and Python/ML deployment pipelines for high‑velocity releases..

Successful candidates bring Experience Handling 500+ Machine Deployments, Experience Scaling Mission-Critical Deployments, and Deep Experience With Kubernetes. Important skills include Command Line Fluency, Keyboard Shortcuts Mastery, Ownership Of Complex Systems, Automation, Problem Solving, and Kubernetes.

Skills & qualifications

RequiredNice to have

Skills

Command Line FluencyKeyboard Shortcuts MasteryOwnership of Complex SystemsAutomationProblem SolvingKubernetesHelmCI/CD OptimizationTypeScript Deployment PipelinesPython/ML Deployment PipelinesDeveloper Experience SupportTerraformPostgreSQLRedisKafka

Qualifications

Experience Handling 500+ Machine DeploymentsExperience Scaling Mission-Critical DeploymentsDeep Experience With KubernetesDeep Experience With HelmDeep Experience With TerraformAbility to Architect and Maintain Complex, Distributed Systems With High-Availability RequirementsHands-on Experience Optimizing Deployment Pipelines for Application Code (TypeScript)Hands-on Experience Optimizing Deployment Pipelines for Machine Learning Models (Python/ML)Excitement About Working Five Days Per Week in San Francisco Office

Full job description

SRE

Location: San Francisco, CA (5 Days In-Office)

You are the infrastructure expert who enables our rapid product development and guarantees 99.9%+ stability and performance of our clinical AI platform for major health systems. Your focus on operational excellence is directly tied to a patient's access to life-saving treatment.

WHAT WE LOOK FOR IN A GREAT ENGINEER

You have the intensity and technical mastery to own mission-critical infrastructure. You hold yourself and others to high standards and thrive in a high-energy, in-office culture where everyone is in it to win it.

  • Tool Proficiency: You are highly proficient with your tools—you speak command line fluently and have mastered keyboard shortcuts.

  • Ownership: You thrive on owning complex systems and have a proven track record of scaling mission-critical deployments.

  • Automation Drive: You love automating things, always finding new ways to increase your own leverage, and defining standards for operational excellence.

  • Problem Solver: You won't wait for someone else to solve a problem that you're in a position to solve; you are willing to jump into whatever needs to get done.

WHAT YOU'LL WORK ON (RESPONSIBILITIES)

As our SRE, you will own the entire production environment and improve the development experience:

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.

  • Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.

  • CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.

  • DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.

  • Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.

TECHNICAL QUALIFICATIONS & ENVIRONMENT

  • IaC & Orchestration: Deep, demonstrable experience with Kubernetes, Helm, and Terraform.

  • Scaling Systems: Proven ability to architect and maintain complex, distributed systems with high-availability requirements.

  • Deployment Experience: Hands-on experience optimizing deployment pipelines for both application code (TypeScript) and machine learning models (Python/ML). Also PostgreSQL, Redis, Kakfa.

  • Core Team Member: Excitement about working five days per week in our San Francisco office.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.