Senior/Staff Cloud Reliability Engineer
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Requirements
Credentials this posting asks for.
Job overview
ThoughtSpot seeks a Staff Cloud Reliability Engineer to own availability, reliability, security, and efficiency of a multi‑cloud SaaS platform. The role demands hands‑on operation of large‑scale Kubernetes control and data planes, AI‑augmented automation, capacity planning, and cloud‑native database management across AWS and GCP environments.
Skills & qualifications
Skills
Qualifications
Full job description
Staff Cloud Reliability Engineer
We are seeking a Staff Site Reliability Engineer with deep enterprise SaaS operations expertise to own the availability, reliability, security, and efficiency of our Multi-Cloud (AWS, GCP) production SaaS platform. The ideal candidate brings hands-on experience running highly available, large-scale Kubernetes-based control and data planes, a strong bias toward automation and AI-augmented operations, and a proven track record in production security, capacity management, and cloud-native data infrastructure. Responsibilities
- Operate a high-scale, multi-cloud (AWS, GCP) SaaS platform — ensuring reliability, performance, and uptime for business-critical production workloads.
- Embed AI and Agentic workflows into SRE practice: leverage AI Ops platforms and LLM-powered autonomous agents for anomaly detection, automated triage, runbook execution, and incident summarization to reduce MTTR.
- Drive capacity planning and scaling operations — proactively model growth, right-size infrastructure, and implement horizontal/vertical autoscaling strategies to support SaaS growth without reliability regression.
- Architect and operate Kubernetes controller frameworks governing both control plane and data plane services; define and enforce operational standards for cluster lifecycle, workload scheduling, autoscaling, and failover.
- Own operations of high-scale cloud-native databases and data infrastructure: PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/OpenSearch, ElastiCache (Redis/Memcached) on AWS and GCP — including performance tuning, backup/recovery, and incident response.
- Lead incident response and blameless post-mortems for P0/P1 events; drive root cause analysis to permanent resolution and prevention — eliminating repeat incidents through systemic fixes, not workarounds.
- Define and enforce a culture of automation-first: identify and eliminate toil through self-healing systems, automated remediation pipelines, and infrastructure-as-code (Terraform, Helm, GitOps).
- Participate in on-call rotations for critical cloud infrastructure; serve as a senior escalation point and incident commander during high-severity events.
- Achieve quantifiable SaaS operational Excellence measured by related SLI/SLO/SLA
Required skills/qualifications
- B.Tech. degree in Computer Science or equivalent.
- At least 6+ years of Enterprise SaaS Ops experience
- Strong proficiency in programming, particularly with Go and Python, and experience with Infrastructure as Code (IaC) tools like Terraform and Ansible.
- Expertise in Cloud Security and/or Cloud networking
- Experience with AI Ops tools, Agentic LLM.
- Experience/ Knowledge in Cloud Services, Kubernetes, Cloud Databases like Postgres/RDS/MySQL/DynamoDB, Elastic, Kafka, and Microservice architecture is a bonus.
- Experience in implementing and operating enterprise-grade observability ( metrics, logs, tracing), alerting stack in a Cloud SaaS environment
- Strong debugging and problem-solving skills (network, systems, database, and application).
- Advanced professional certifications from Cloud Providers ( AWS, Azure, GCP) in domains like K8s, Solution architecture, networking, and databases are a bonus.
- Full Stack Architecture/Development Experience is a bonus.
You've read the whole posting — now see how you match it.