Senior Site Reliability Engineer
United StatesJobPosted 3 days agoStill listed 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Laravel seeks a Senior Site Reliability Engineer to design, build, and maintain multi‑region Kubernetes infrastructure, create observability systems with Prometheus and Grafana, and automate operations. The remote role collaborates across product, security, and engineering teams, establishing SLOs/SLIs, driving reliability culture, and contributing to Laravel Cloud, Nightwatch, and Forge services.
Skills & qualifications
Skills
Benefits
Full job description
At Laravel, we don’t just build tools; we build the foundation that empowers millions of developers to ship their dreams. We are looking for a Senior Site Reliability Engineer to help us scale that mission by ensuring our global infrastructure remains as elegant and reliable as the code we write. If you are energized by the challenge of building robust observability systems, running Kubernetes clusters across multiple regions, and solving complex operational puzzles with code, you’ve found your next home. Location / Timezone: Remote, between Central Europe and US East timezones for optimal collaboration with the team. Description of the Role As a Senior Site Reliability Engineer, you will be an early contributor to our dedicated SRE function, reporting directly to Florian Beer. This is a high-impact, autonomous role where you will design and implement the systems that support Laravel Cloud, Nightwatch, and Forge. You will act as a bridge between development and operations, advocating for a blameless culture and shared responsibility for reliability across the entire organization. Your 12-Month Mission Imagine we are all at Laracon in 12 months' time. You are telling the team about your first year, and the impact is undeniable:
- First 30 Days: You will have selected an SLO monitoring solution.
- Day 60: You will have guided teams towards identifying SLIs and how/why we work with SLOs at Laravel.
- Day 90: You have established clear, data-driven SLOs across engineering teams, giving us a unified language for reliability.
- Year One: You have worked together with our engineering teams to identify SLIs, review existing SLAs, created SLOs, and educated teams on how SLO’s error budgets guide decisions.
What You Will Do
- Architect Reliability: Establish SRE as a core function at Laravel, building the fundamentals from the ground up.
- System Design: Design, build, and maintain multi-region Kubernetes infrastructure and global distributed systems.
- Automation: Solve operational challenges through software, reducing manual intervention (toil) for our product teams.
- Observability: Design and implement monitoring, logging, and alerting systems using tools like Prometheus, Grafana, and Loki.
- Collaboration: Partner with product leads and SecOps to make reliability a shared responsibility.
Requirements - What You Will Bring
- Infrastructure Mastery: Deep experience with Linux system administration and cloud platforms, specifically AWS.
- Orchestration & IaC: Proficiency with Kubernetes, Docker, and managing infrastructure via Terraform.
- Programming Skills: The ability to solve problems with software and scripting using e.g. PHP, Bash, or Go.
- Systems Thinking: A "smart and passionate" approach to troubleshooting, with the ability to deconstruct complex systems into triagable components.
- Reliability Mindset: Experience with SLO/SLI/SLA definition, capacity planning, and performance tuning.
- Soft Skills: A commitment to documentation, cross-team collaboration, and an automation-first mindset.
Requirements - Bonus Skills
-
Framework Familiarity: Previous experience working with the Laravel framework and our existing product suite (Cloud, Forge, Vapor, etc.) is highly preferred.
-
Advanced Observability: Experience with Prometheus, Grafana Mimir, and Grafana Loki for metrics storage and alerting.
-
Small tight-knit team where every developer counts
-
Fully remote and globally distributed working environment
-
Option to attend Laracon conferences around the world
-
Health care plan (Medical, Dental & Vision)
-
Paid time off (Vacation, Sick & Public holidays)
-
Family leave (Maternity, Paternity)
-
Pension plans (As locally applicable)
-
Performance based bonus plan
-
Company equity
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Site Reliability EngineerReplit · Remote · US · $210–275K/yrPosted 2w agoPosted 2w ago
Senior Site Reliability EngineerDocebo Inc. · Remote · US · $145–193K/yrPosted 3w agoPosted 3w agoSenior SRE Banyan Software · Remote · US · $145–170K/yrPosted 1w agoPosted 1w ago
Senior Site Reliability Engineer (FedRAMP)Cisco · RTP, NC (Hybrid) · $139–204K/yrPosted 3 days agoPosted 3 days agoSenior Site Reliability Engineer, SecurityAuthzed · Remote · USPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.