Cubbit logo

Site Reliability Engineer (SRE)

Cubbit

Remote · location unlistedFull-timePosted 3mo agoStill listed 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Remote · location unlisted
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Cubbit is hiring a Site Reliability Engineer (SRE). Cubbit is seeking a Mid-Level Site Reliability Engineer to join its Tech Operations team. The role involves ensuring the platform's health, reliability, and performance, automating operations, and improving observability. The engineer will collaborate with other teams to solve production challenges and build resilient systems.

Key focus areas include Keep Cubbit's geo-distributed platform healthy, reliable, and performant across production environments, Build automation that eliminates manual work and makes operations safer and faster, and Improve observability through meaningful metrics, dashboards, logging, and alerting.

Successful candidates bring 3+ Years Site Reliability Engineering Experience and Professional Proficiency In English. Important skills include Linux, Kubernetes, Containerized Workloads, Infrastructure as Code, Networking Fundamentals, and DNS. Preferred (not required): Elastic Stack, Technical Support L2/L3, Technical Operations, and Customer-Facing Infrastructure Support.

Skills & qualifications

RequiredNice to have

Skills

LinuxKubernetesContainerized WorkloadsInfrastructure as CodeNetworking FundamentalsDNSTCP/IPHTTPTLSLoad BalancingMonitoring ToolsObservability ToolsBashPythonCI/CD PipelinesModern Deployment PracticesTroubleshooting Production SystemsEnglishPulumiPrometheusGrafanaLokiElastic StackOpenSearchTechnical Support L2/L3Technical OperationsCustomer-Facing Infrastructure SupportManaging Complex Production IncidentsTechnical EscalationsDistributed SystemsCloud InfrastructureS3-Compatible Object StorageGitOps PracticesArgo CDFluxAWSAzureGCPSecurity Best PracticesInfrastructure HardeningOpen-Source ContributionSystems ThinkingAutomationOwnershipRoot Cause AnalysisCollaborationCalm Under PressureCuriosityEagerness to LearnPragmatism

Qualifications

3+ Years Site Reliability Engineering Experience3+ Years DevOps Experience3+ Years Platform Engineering Experience

Full job description

We not only apply cutting-edge technology. We create it. At Cubbit, we're building the next generation of cloud storage: globally distributed, geo-distributed by design, and independent from hyperscalers. Reliability is at the heart of everything we do.

We're looking for a Mid-Level Site Reliability Engineer to join our Tech Operations team and help keep our platform resilient, scalable, and always available. You'll work at the intersection of software and infrastructure , collaborating with engineering teams to automate operations, improve observability, and solve complex production challenges before our users even notice them.

If you enjoy building reliable systems, automating everything that shouldn't be done twice, and turning operational pain into engineering improvements, we'd love to meet you.

‍

What you’ll do

  • Keep Cubbit's geo-distributed platform healthy, reliable, and performant across production environments.
  • Build automation that eliminates manual work and makes operations safer and faster.
  • Improve observability through meaningful metrics, dashboards, logging, and alerting.
  • Investigate production incidents and complex customer escalations, perform root cause analysis, and turn learnings into long-term reliability improvements.
  • Partner with Software Engineers to design systems that are reliable by default.
  • Continuously improve CI/CD pipelines and deployment workflows.
  • Participate in the on-call rotation while helping make it quieter every week.
  • Optimize performance, scalability, and operational efficiency across our infrastructure.
  • Contribute to operational best practices, documentation, and a healthy engineering culture. ‍

What you’ll need

  • 3+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar roles.
  • Strong Linux administration skills and a solid understanding of production environments.
  • Hands-on experience with Kubernetes and containerized workloads.
  • Experience with Infrastructure as Code tools (Pulumi preferred).
  • Good understanding of networking fundamentals (DNS, TCP/IP, HTTP, TLS, load balancing).
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar.
  • Scripting skills in Bash, Python, or similar languages.
  • Familiarity with CI/CD pipelines and modern deployment practices.
  • Confidence troubleshooting production systems under pressure.
  • Professional proficiency in English. ‍

Bonus point You'll stand out if you have experience with:

  • Technical Support (L2/L3), Technical Operations, or customer-facing infrastructure support.
  • Managing complex production incidents and technical escalations.
  • Distributed systems, cloud infrastructure, or S3-compatible object storage.
  • GitOps practices and tools such as ArgoCD or Flux.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Security best practices, infrastructure hardening, or open-source contribution. ‍

Profile and mindset We're looking for someone who:

  • Thinks in systems, not just servers.
  • Automates first and hates doing the same thing twice.
  • Takes ownership and follows problems through to the root cause: not just the quick fix.
  • Thrives in collaborative environments where engineering and operations work as one team.
  • Enjoys working in a fast-paced scale-up environment, where priorities evolve, ownership is expected, and everyone contributes beyond their role.
  • Stays calm when production gets noisy and enjoys solving problems that matter.
  • Is curious, eager to learn, and continuously looking for ways to improve.
  • Balances pragmatism with engineering excellence: you know when "good enough" is actually the best solution.
  • Believes that reliability is a feature, not an afterthought. ‍

Why you’ll love working with us

  • Join an ambitious and supportive team of builders and innovators
  • Remote-friendly environment with flexible schedules
  • Opportunity to contribute to one of Europe’s most advanced cloud technologies ‍

Location Full remote work is welcome in Cubbit, and you can freely use the lounge area in one of our co-working partner offices worldwide. But remember that our headquarters, in the centre of Bologna, Italy, is always available for you. Let's find the right workspace for you.

‍

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.