Runlayer logo

Member of Technical Staff - Site Reliability

Runlayer

New York, NYRemoteFull-time$140–250K/yrPosted 4mo agoVerified open 1w ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$140–250K/yr
Location
New York, NYRemote
Schedule
Full-time
Work Authorization
Not specified

Job overview

Runlayer is hiring a Member of Technical Staff - Site Reliability. Runlayer seeks a Site Reliability Engineer to own reliability, performance, and scalability of its AI platform infrastructure, supporting enterprise customers across cloud and on‑prem environments while collaborating closely with founders and a senior engineering team.

Key focus areas include Own reliability and performance of cloud infrastructure across AWS (ECS, Aurora, CloudWatch) and GCP, Manage and optimize Kubernetes clusters and container orchestration, and Drive database reliability engineering, including performance tuning and scaling.

Successful candidates bring Background At A B2B SaaS Company Serving Enterprise Customers. Important skills include AWS, Amazon ECS, Aurora, CloudWatch, GCP, and Kubernetes. Preferred (not required): On-Prem Deployment, Hybrid Environment Support, Python Backend Familiarity, and On Prem Or Hybrid Environments.

Skills & qualifications

RequiredNice to have

Skills

AWSAmazon ECSAuroraCloudWatchGCPKubernetesContainer OrchestrationDBREDatabase Performance TuningCI/CD Pipeline OwnershipIncident ManagementOn-Prem DeploymentHybrid Environment SupportPython Backend FamiliarityCI/CD PipelinesOn Prem or Hybrid EnvironmentsPython Backend

Qualifications

Background at B2B SaaS Company Serving Enterprise CustomersExperience Deploying and Supporting on Prem or Hybrid EnvironmentsExperience at Early Stage or High Growth Company

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
Paid Time Off
Parental Leave

Full job description

About Runlayer AI is transforming how every company operates, but most enterprises are stuck. They want to move fast with AI Agents, tools, and workflows, but they can't do it safely. We're fixing that.

Our team built AI Actions for OpenAI, shipped Zapier Agents to millions of users, and launched the first remote MCP server with Anthropic. We helped establish the protocol, and now we're building the platform enterprises need to actually put AI to work.

Runlayer is one platform for MCPs, Skills, and Agents : purpose-built security, fine-grained governance, and complete observability so organizations can go all-in on AI across the entire company without the risk. We just raised a $30M Series A led by Felicis, with participation from Khosla Ventures, bringing our total raised to $42M. Already trusted by Gusto, Instacart, Opendoor, dbt Labs, and Decagon.

About the Role As our Site Reliability Engineer, you'll own the reliability, performance, and scalability of Runlayer's infrastructure as we grow to serve enterprise customers across cloud and on-prem environments.

Why You'll Thrive Here

  • Impact: Build the infrastructure foundation for the enterprise MCP platform, directly enabling AI adoption at scale

  • Excellence: Work closely with founders and a small, senior engineering team shipping fast in a high-growth environment

  • Ownership: Own reliability end-to-end, from database performance to incident response to CI/CD pipelines

What You'll Do

  • Own reliability and performance of our cloud infrastructure across AWS (ECS, Aurora, CloudWatch) and GCP

  • Manage and optimize Kubernetes clusters and container orchestration

  • Drive database reliability engineering, including performance tuning and scaling

  • Build and maintain CI/CD pipelines for rapid, safe deployments

  • Run incident response and on-call rotations

  • Partner with product engineers to design scalable, resilient systems

What We're Looking For

  • Strong AWS experience, particularly ECS, Aurora, and CloudWatch

  • GCP experience as we expand cross-cloud

  • Kubernetes and container orchestration expertise

  • DBRE experience with database performance tuning

  • CI/CD pipeline ownership and incident response experience

  • Background at a B2B SaaS company serving enterprise customers, ideally in infrastructure

Bonus Qualifications

  • Experience deploying and supporting on-prem or hybrid environments

  • Python backend familiarity (our platform is Python-based)

  • Experience at an early-stage or high-growth company

What We Offer We provide a competitive package designed to attract and retain top talent who can work effectively with enterprise customers.

  • Competitive salary and equity — compensation that reflects your expertise and customer-facing responsibilities.

  • Paid time off — paid vacation, paid sick leave, and paid parental leave.

  • Professional development — budget for conferences, courses, and certifications in AI, enterprise software, and customer success.

  • Top-tier equipment — your choice of laptop and accessories to create your ideal work environment.

  • Health benefits — comprehensive health, dental, and vision coverage.

  • Customer interaction opportunities — work directly with innovative companies and see the immediate impact of your work.

Not quite the right fit? Reach out to [email protected] with details about your experience and interests.

You've read the whole posting — now see how you match it.