Senior DevOps Engineer
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
The Senior DevOps Engineer will design, implement, and operate scalable AWS infrastructure, own end‑to‑end production systems, and drive reliability, automation, and cost‑optimization across engineering and data teams.
Skills & qualifications
Skills
Qualifications
Full job description
We are looking for an engineer who has designed and owned production infrastructure end-to-end, rather than someone who has only maintained existing DevOps tooling. The ideal candidate combines deep AWS and Terraform expertise with strong CI/CD ownership, reliability engineering, automation, and the ability to work across engineering and data teams.
Responsibilities:
- Architect, implement, and operate scalable AWS infrastructure supporting business-critical applications and data platforms.
- Design and manage Infrastructure as Code (IaC) using Terraform.
- Build reusable Terraform modules and establish best practices around state management, environments, and infrastructure lifecycle.
- Modernize existing infrastructure by importing and migrating brownfield AWS resources into Terraform.
- Design, build, and own CI/CD pipelines from code commit through production deployment.
- Implement progressive delivery strategies such as blue-green and canary deployments.
- Build automated rollback mechanisms based on application health, deployment metrics, and SLO breaches.
- Manage and optimize AWS services, including EC2 RDS, ElastiCache, IAM, VPC, networking, and cost-management tooling.
- Build and operate stateful systems supporting real-time and batch analytics at a large scale.
- Partner with engineering, data, analytics, and product teams to translate business requirements into production-grade infrastructure.
- Establish and improve observability across infrastructure and applications using metrics, logs, traces, and alerts.
- Define and monitor Service Level Objectives (SLOs) and drive improvements in platform reliability.
- Automate operational workflows and reduce manual intervention and on-call burden.
- Troubleshoot complex infrastructure, networking, deployment, and production reliability issues.
- Drive cloud cost optimization and infrastructure efficiency across AWS environments.
- Establish engineering standards around security, scalability, reliability, and operational excellence.
Requirements:
- 7+ years of experience in DevOps, platform engineering, SRE, or infrastructure engineering.
- Strong hands-on experience with AWS in production environments.
- Deep knowledge of EC2 RDS, ElastiCache, IAM, VPC, AWS networking, and cloud cost optimization.
- Strong expertise in Terraform with experience in Terraform module architecture, state management, remote state, infrastructure lifecycle management, brownfield infrastructure import and migration, and multi-environment infrastructure.
- Proven experience owning production CI/CD pipelines.
- Hands-on experience with GitHub Actions or equivalent CI/CD platforms.
- Experience implementing progressive delivery and automated deployment rollback.
- Strong understanding of cloud networking, security, IAM, and infrastructure architecture.
- Experience with observability and monitoring platforms.
- Strong understanding of reliability engineering, SLOs, incident management, and production operations.
- Experience working with large-scale data, analytics, or distributed systems environments.
Good to Have:
- Experience with Kubernetes and containerized workloads.
- Experience with AWS CloudFormation or other IaC technologies.
- Experience with ArgoCD, Spinnaker, GitLab CI, Jenkins, or similar deployment platforms.
- Experience with Prometheus, Grafana, Datadog, CloudWatch, OpenTelemetry, or similar observability tools.
- Experience operating infrastructure supporting large-scale or petabyte-scale data platforms.
- Strong scripting skills in Python, Bash, or similar languages.
- Experience with FinOps and cloud cost optimization.
- Experience designing highly available and fault-tolerant systems.
Core Tech Stack:
- Cloud: AWS.
- IaC: Terraform.
- CI/CD: GitHub Actions / equivalent.
- Infrastructure: EC2 RDS, ElastiCache, VPC, IAM.
- Observability: CloudWatch / Prometheus / Grafana / Datadog / OpenTelemetry.
- Scripting: Python / Bash.
- Deployment: Progressive Delivery, Automated Rollbacks.
You've read the whole posting — now see how you match it.