Beam logo

Site Reliability Engineer

Beam

New York, NYFull-timeSeen 1mo agoStill listed 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
New York, NY
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Beam is an ultrafast AI inference platform seeking a Site Reliability Engineer to own compute fleet health, build metrics pipelines and alerting, automate deployment debugging, design GPU qualification processes, and manage firmware‑level telemetry and log collection at scale.

Skills & qualifications

RequiredNice to have

Skills

Hardware InstinctAI ToolingProduction DebuggingDeveloper Tools EnthusiasmCloud Native TechnologiesOpen Source Software

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
Tuition Assistance

Full job description

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub. About the Role

  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production.
  • Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage.
  • Design the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production.
  • Own firmware-level telemetry, log collection at scale, and the low-level access layer that repair automation and health tooling depend on.

Skills & Experience

  • You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it
  • You're fluent with AI tooling. You aren’t afraid to max-out your token usage for the right spec.
  • You’re comfortable debugging production issues, from triage to post-mortem.
  • Enthusiasm for developer tools, cloud native technologies, and open source software

Benefits

  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.