
Site Reliability Engineer
Berkeley, CA · HybridContract$80/hrSeen 1mo agoSeen in employer's feed 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Berkeley, CA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
The national HPC facility seeks a sharp, self‑motivated Site Reliability Engineer to monitor, automate, and improve compute, storage, network, and facility systems, ensuring uninterrupted service for thousands of scientists. The role involves real‑time alert triage, tool development, data‑center floor operations, and coordination of maintenance activities across teams.
Skills & qualifications
Skills
Full job description
📍 Hybrid — Berkeley, CA
📅 1 Year Contract Assignment with possibility of extension based on performance and organizational needs.
💰 $80/hr
Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.
If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.
What You'll Own
-
Monitor and triage alerts across compute, storage, network, and facility systems in real time
-
Build automation that prevents issues before they become outages
-
Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
-
Walk the data center floor to keep power, cooling, and environmental systems humming
-
Coordinate maintenance activities across teams and keep incidents accurately tracked
-
Dig into complex, ambiguous problems and drive them to resolution
What You Bring
-
Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
-
Solid Linux/command-line (SSH) chops
-
Programming/scripting experience: Python, C, C++, Perl, or Java
-
A self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
-
Network security fundamentals (ACLs, firewalls)
-
Strong cross-team communication and collaboration skills
Nice to Have
-
Experience building or deploying Agentic AI / autonomous automation for technical workflows
-
ServiceNow implementation experience
-
ITSM best-practice know-how
Powered by JazzHR
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Staff Site Reliability EngineerFivetran · Oakland, CA (Hybrid) · $232–290K/yrPosted 4w agoPosted 4w ago
Engineering Manager, Site Reliability EngineeringReplit · Foster City, CA (Hybrid) · $250–325K/yrPosted 2w agoPosted 2w ago
Senior/Lead Software Engineer, Site Reliability (Agentforce Operations)Salesforce · San Francisco, CA · $149–314K/yrPosted 1w agoPosted 1w agoForward Deployed Engineer - PlatformDitto · San Diego, CA (Hybrid) · $156–189K/yrPosted todayPosted today
- Senior Systems Engineer - Enterprise AI PlatformsLambda · San Francisco, CA (Hybrid) · $206–275K/yrPosted 2 days agoPosted 2 days ago
You've read the whole posting — now see how you match it.