Groww logo

Site Reliability Engineer

Groww

Bangalore, Karnataka, IndiaJobPosted 2mo agoStill listed today

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Bangalore, Karnataka, India
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Requirements

Credentials this posting asks for.

Bachelor's degree

Job overview

The Site Reliability Engineer will design scalable, reliable infrastructure, implement high‑availability strategies, define SLOs/SLIs, lead incident management, mentor junior engineers, build monitoring systems, evolve IaC practices, and manage large Kubernetes clusters.

Skills & qualifications

RequiredNice to have

Skills

Infrastructure as Code (IaC)DatadogMentorshipAmazon Web ServicesLoad BalancingMicrosoft AzureKubernetesIncident ManagementScriptingSystem ArchitectureELKPrometheusDevOpsOrchestrationDNSContainersPythonTerraformGoogle Cloud PlatformSplunkGolangSOA/MicroservicesNetworking TechnologiesJavaAnsibleSREGoIaCELK StackNew RelicMonitoringNetworkingAWSAzureGCPLeadershipCommunicationCritical ThinkingMicroservicesContainer Orchestration

Qualifications

Bachelor's/Master's Degree in Computer Science or Engineering or Equivalent Experience6-9 Years in SRE, DevOps, or System Architecture Roles

Full job description

Responsibilities:

  • Architect and lead the design of scalable, reliable infrastructure solutions.
  • Implement strategies for high availability, scalability, and low-latency performance.
  • Define service-level objectives (SLOs) and indicators (SLIs) to track performance and reliability.
  • Drive incident management, identifying root causes and providing long-term solutions.
  • Mentor junior engineers and foster a collaborative, learning-focused environment.
  • Design advanced monitoring and alerting systems for proactive system management.
  • Evolve Infrastructure as Code (IaC) practices to automate infrastructure provisioning.
  • Collaborate on reliability roadmaps, performance benchmarks, and disaster recovery plans.
  • Manage Kubernetes clusters at scale, integrating service meshes like Istio or Linkerd.
  • Implement chaos engineering principles for system resilience.
  • Influence technical direction, reliability culture, and organizational strategies.

Requirements:

  • Mandatory Skills: SRE, Python, Go, Iac, Datadog, Terraform, Ansible, Splunk, Prometheus, ELK Stack, New Relic, Monitoring, Networking, Load Balancing, Incident Management, DevOps, Kubernetes, AWS or Azure or GCP, Bachelor's/Master's degree in Computer Science or Engineering, or equivalent experience.
  • 6-9 years in SRE, DevOps, or system architecture roles with large-scale production systems.
  • Proven experience in managing complex cloud environments (AWS, GCP, Azure).
  • Expertise in Kubernetes, container orchestration, and microservices.
  • Advanced programming/scripting skills in Python, Go, or Java.
  • Proficiency in monitoring/logging tools (Prometheus, ELK Stack, New Relic).
  • Strong understanding of networking, load balancing, and DNS management.
  • Leadership skills with experience mentoring teams and collaborating with stakeholders.
  • Strong communication, critical thinking, and incident management abilities.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.