
Senior Staff SRE – Compute Platform
Bengaluru, Karnataka, IndiaFull-timePosted 2w agoStill listed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Bengaluru, Karnataka, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
NVIDIA seeks a Senior Staff Site Reliability Engineer to build and operate reliable, scalable compute platforms supporting global engineering workloads, focusing on Kubernetes, KubeVirt, bare‑metal infrastructure, automation, observability, and AI‑enabled operations.
Skills & qualifications
Skills
Qualifications
Full job description
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services.
What you’ll be doing:
-
Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale.
-
Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.
-
Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data.
-
Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems.
-
Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation.
What we need to see:
-
BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services.
-
Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges.
-
Experience deploying and operating bare-metal infrastructure in a data-center environment, including provisioning, networking, operating-system lifecycle management, and hardware automation.
-
Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.
-
Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet, along with a solid understanding of TCP/IP networking and infrastructure security.
-
Strong SRE and observability experience, including SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.
-
Clear written and interpersonal communication skills, with a record of delivering practical, scalable solutions to complex technical problems.
Ways to stand out from the crowd:
-
Experience operating HPC, AI, GPU-accelerated, or general-purpose bare-metal compute infrastructure, including GPU-enabled Kubernetes or KubeVirt clusters.
-
Expertise with VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.
-
Experience applying generative AI or agentic workflows to improve infrastructure diagnostics, reduce operational toil, and accelerate incident resolution.
-
Experience building secure, integrated operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, workflow orchestration, and infrastructure or incident-management systems.
-
Demonstrated delivery of complex, high-impact infrastructure projects.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Staff Site Reliability EngineerNVIDIA · Bengaluru, India, IndiaPosted 5 days agoPosted 5 days ago
Senior Staff Site Reliability EngineerNVIDIA · Bengaluru, Karnataka, India (Hybrid)Posted 1w agoPosted 1w ago
Senior Site Reliability EngineerCisco · Bangalore, Karnataka, India (Hybrid)Posted 2w agoPosted 2w ago
Site Reliability EngineerCisco · Bangalore, Karnataka, India (Hybrid)Posted todayPosted today
Staff Site Reliability Engineer - Monitoring and Anomaly Detection (Monetization) GitLab · Remote · INPosted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.