Site Reliability Engineer (SRE)
Remote · location unlistedFull-timePosted 3mo agoStill listed 4 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New remote roles like this one, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Cubbit is hiring a Site Reliability Engineer (SRE). Cubbit is seeking a Mid-Level Site Reliability Engineer to join its Tech Operations team. The role involves ensuring the platform's health, reliability, and performance, automating operations, and improving observability. The engineer will collaborate with other teams to solve production challenges and build resilient systems.
Key focus areas include Keep Cubbit's geo-distributed platform healthy, reliable, and performant across production environments, Build automation that eliminates manual work and makes operations safer and faster, and Improve observability through meaningful metrics, dashboards, logging, and alerting.
Successful candidates bring 3+ Years Site Reliability Engineering Experience and Professional Proficiency In English. Important skills include Linux, Kubernetes, Containerized Workloads, Infrastructure as Code, Networking Fundamentals, and DNS. Preferred (not required): Elastic Stack, Technical Support L2/L3, Technical Operations, and Customer-Facing Infrastructure Support.
Skills & qualifications
Skills
Qualifications
Full job description
We not only apply cutting-edge technology. We create it. At Cubbit, we're building the next generation of cloud storage: globally distributed, geo-distributed by design, and independent from hyperscalers. Reliability is at the heart of everything we do.
We're looking for a Mid-Level Site Reliability Engineer to join our Tech Operations team and help keep our platform resilient, scalable, and always available. You'll work at the intersection of software and infrastructure , collaborating with engineering teams to automate operations, improve observability, and solve complex production challenges before our users even notice them.
If you enjoy building reliable systems, automating everything that shouldn't be done twice, and turning operational pain into engineering improvements, we'd love to meet you.
What you’ll do
- Keep Cubbit's geo-distributed platform healthy, reliable, and performant across production environments.
- Build automation that eliminates manual work and makes operations safer and faster.
- Improve observability through meaningful metrics, dashboards, logging, and alerting.
- Investigate production incidents and complex customer escalations, perform root cause analysis, and turn learnings into long-term reliability improvements.
- Partner with Software Engineers to design systems that are reliable by default.
- Continuously improve CI/CD pipelines and deployment workflows.
- Participate in the on-call rotation while helping make it quieter every week.
- Optimize performance, scalability, and operational efficiency across our infrastructure.
- Contribute to operational best practices, documentation, and a healthy engineering culture.
What you’ll need
- 3+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar roles.
- Strong Linux administration skills and a solid understanding of production environments.
- Hands-on experience with Kubernetes and containerized workloads.
- Experience with Infrastructure as Code tools (Pulumi preferred).
- Good understanding of networking fundamentals (DNS, TCP/IP, HTTP, TLS, load balancing).
- Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar.
- Scripting skills in Bash, Python, or similar languages.
- Familiarity with CI/CD pipelines and modern deployment practices.
- Confidence troubleshooting production systems under pressure.
- Professional proficiency in English.
Bonus point You'll stand out if you have experience with:
- Technical Support (L2/L3), Technical Operations, or customer-facing infrastructure support.
- Managing complex production incidents and technical escalations.
- Distributed systems, cloud infrastructure, or S3-compatible object storage.
- GitOps practices and tools such as ArgoCD or Flux.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Security best practices, infrastructure hardening, or open-source contribution.
Profile and mindset We're looking for someone who:
- Thinks in systems, not just servers.
- Automates first and hates doing the same thing twice.
- Takes ownership and follows problems through to the root cause: not just the quick fix.
- Thrives in collaborative environments where engineering and operations work as one team.
- Enjoys working in a fast-paced scale-up environment, where priorities evolve, ownership is expected, and everyone contributes beyond their role.
- Stays calm when production gets noisy and enjoys solving problems that matter.
- Is curious, eager to learn, and continuously looking for ways to improve.
- Balances pragmatism with engineering excellence: you know when "good enough" is actually the best solution.
- Believes that reliability is a feature, not an afterthought.
Why you’ll love working with us
- Join an ambitious and supportive team of builders and innovators
- Remote-friendly environment with flexible schedules
- Opportunity to contribute to one of Europe’s most advanced cloud technologies
Location Full remote work is welcome in Cubbit, and you can freely use the lounge area in one of our co-working partner offices worldwide. But remember that our headquarters, in the centre of Bologna, Italy, is always available for you. Let's find the right workspace for you.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Site Reliability EngineerMaintainX · Remote · US · $120–249K/yrPosted 3w agoPosted 3w ago
Site Reliability Engineer IIIVida Health · Remote · US · $175–185K/yrPosted 2 days agoPosted 2 days agoStaff Site Reliability EngineerReplit · Remote · US · $250–325K/yrPosted 2w agoPosted 2w ago
Lead Site Reliability EngineerPeraton · Remote · US · $112–179K/yrPosted 3 days agoPosted 3 days ago
Senior Site Reliability EngineerDocebo Inc. · Remote · US · $145–193K/yrPosted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.