
Senior Site Reliability Engineer
Salt Lake City, UT · HybridJobSeen 2mo agoSeen in employer's feed 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Salt Lake City, UT, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
PrincePerelson & Associates is hiring a Senior Site Reliability Engineer. The Senior Site Reliability Engineer will build highly reliable, scalable cloud platforms for mission-critical applications. This role involves shaping reliability standards, improving developer experience, and building resilient systems. The engineer will work at the intersection of software engineering and infrastructure, partnering with teams to improve system performance, automate operations, and establish best practices across Kubernetes-based services.
Key focus areas include Define and evolve Service Level Objectives (SLOs) and Service Level Indicators (SLIs)., Build and standardize monitoring, alerting, and observability practices across engineering teams., and Develop scalable observability solutions leveraging metrics, logs, traces, and distributed tracing technologies..
Successful candidates bring 7+ Years Site Reliability Engineering Experience, 7+ Years Platform Engineering Experience, and 7+ Years DevOps Experience. Important skills include Terraform, OpenTelemetry, Kubernetes, SLOs, SLIs, and Monitoring. Preferred (not required): PostgreSQL, MongoDB, DynamoDB, and Prometheus.
Skills & qualifications
Skills
Qualifications
Benefits
Full job description
Senior Site Reliability Engineer (SRE)
Salt Lake City, UT
Are you passionate about building highly reliable, scalable cloud platforms that power mission-critical applications? We're partnering with an innovative technology company that's investing heavily in platform reliability, automation, and observability. This is an opportunity to have a significant impact on the engineering organization by shaping reliability standards, improving developer experience, and helping build resilient systems that support a rapidly growing platform.
You'll work at the intersection of software engineering and infrastructure, partnering with engineering teams to improve system performance, automate operations, and establish best practices across Kubernetes-based services. If you enjoy solving complex distributed systems challenges and creating tools that make engineers more productive, we'd love to talk.
What You'll Do
-
Define and evolve Service Level Objectives (SLOs) and Service Level Indicators (SLIs) that align platform performance with customer and business expectations.
-
Build and standardize monitoring, alerting, and observability practices across engineering teams using Infrastructure as Code (Terraform preferred).
-
Develop scalable observability solutions leveraging metrics, logs, traces, and distributed tracing technologies such as OpenTelemetry.
-
Evaluate and implement modern observability platforms and integrations, leading proof-of-concepts and defining adoption strategies.
-
Establish reliability standards for Kubernetes-based applications, including scaling strategies, deployment safety, resource optimization, dashboards, and alerting.
-
Design automation that reduces manual effort, streamlines operations, and improves incident response and recovery.
-
Lead high-severity incident response efforts, facilitate postmortems, and drive long-term reliability improvements through measurable action plans.
-
Participate in an on-call rotation while continually improving monitoring quality and reducing unnecessary alert noise.
-
Partner with software engineers and platform teams to build resilient, scalable cloud infrastructure and operational best practices.
What We're Looking For
-
7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a similar infrastructure-focused engineering role.
-
Strong experience supporting production workloads in AWS and Kubernetes environments.
-
Demonstrated success implementing SLOs, SLIs, and reliability engineering best practices.
-
Experience designing observability solutions that provide actionable insights while minimizing operational noise.
-
Strong Infrastructure as Code experience, preferably with Terraform.
-
Experience automating operational workflows and building reusable engineering standards.
-
Familiarity with distributed systems and modern data platforms including technologies such as PostgreSQL, MongoDB, DynamoDB, or similar.
-
Experience with multiple observability platforms such as Prometheus, Grafana, New Relic, Splunk, CloudWatch, ELK, or comparable technologies.
-
Strong troubleshooting and debugging skills across distributed production environments.
-
Experience leading incident response and driving continuous operational improvements.
-
Excellent communication skills with the ability to influence engineering teams through collaboration, technical leadership, and practical solutions.
Why This Opportunity?
-
Join a high-performing engineering organization where reliability is a strategic priority.
-
Work with modern cloud-native technologies including AWS, Kubernetes, Terraform, and OpenTelemetry.
-
Influence engineering standards and platform architecture across multiple teams.
-
Solve complex technical challenges at scale while helping shape the future of the platform.
-
Collaborative hybrid environment with four days in the office and one remote day each week.
-
Comprehensive benefits package including medical, dental, vision, retirement plan, generous paid time off, parental leave, and additional wellness benefits.
PrincePerelson & Associates is an Equal Opportunity Employer and complies with all provisions of the EEO and ADA laws. We do not discriminate in our employment practices on the basis of race, color, religion, national origin, sex (including sexual orientation and sexual identity), age, genetic information, parental status, military status, disability, or any non-merit-based factors or other federal, state, or locally protected class. All applicants applying for U.S. job openings must be authorized to work in the United States.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Site Reliability Engineer, Cloud InfrastructureWeave Communications Inc. · Remote · USPosted 1 day agoPosted 1 day ago
Site Reliability Engineer, Cloud InfrastructureWeave · Remote · USPosted 1 day agoPosted 1 day agoSenior Data Insights SpecialistCVS Health · Remote · US · $47–112K/yrPosted 2w agoPosted 2w ago
Senior Data AnalystCadmus · Salt Lake City, UT · $80–95K/yrPosted 1w agoPosted 1w ago
Senior Exercise PlannerCadmus · Salt Lake City, UT · $85–115K/yrPosted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.