Senior Site Reliability engineer ›
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
Filevine is hiring a Senior Site Reliability engineer ›. Filevine seeks a Senior Site Reliability Engineer to design monitoring, logging, and tracing systems, build automation and CI/CD tools, drive reliability initiatives, own incident response, mentor engineers, and apply AI/ML to operational data, ensuring scalable, secure, and resilient production environments.
Key focus areas include Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production visibility, Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and reliability, and Drive implementation and continuous improvement of reliable systems for building, deploying, testing, and operating products.
Important skills include Site Reliability Engineering, Cloud Infrastructure, DevOps, Kubernetes, Automation, and Incident Response.
Skills & qualifications
Skills
Qualifications
Full job description
Responsibilities • Design and improve the monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs that give teams meaningful visibility into production health and customer impact.
-
Build and maintain automation, internal tools, and CI/CD systems that increase engineering efficiency, reduce toil, and support reliable deployments at scale. Take responsibility for the quality and reliability of tools and services you support.
-
Drive the implementation and continuous improvement of reliable systems for building, deploying, testing, and operating Filevine products, proactively identifying and resolving reliability, performance, scalability, and security risks before they impact customers.
-
Own complex production incidents through detection, triage, communication, resolution, and follow-up. Turn incident learning into durable corrective actions, stronger runbooks and operating practices, and improvements that reduce recurring incidents and operational burden.
Lead significant technical initiatives from problem definition and design through implementation and adoption. Coordinate work across engineers and teams, communicate tradeoffs and risks, and help ensure the work delivers the intended results.
-
Mentor other Site Reliability Engineers through design reviews, incident follow-ups, paired problem-solving, and meaningful delegation. Help engineers develop stronger technical judgment and become increasingly capable of handling complex production work independently.
-
Participate in the shared on-call rotation and help ensure production systems are prepared to operate reliably at scale through capacity planning, operational readiness, and continuous improvements to resilience and recovery.
-
Apply AI and machine learning to analyze operational signals, identify patterns, forecast reliability and capacity risks, and implement improvements that make systems more reliable, efficient, and easier to operate.
Qualifications • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles, including at least 5 years in a Site Reliability Engineering or reliability-focused role.
-
Strong knowledge of distributed systems and hands-on experience operating Kubernetes workloads and cloud infrastructure in AWS or a comparable platform, with proficiency in Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.
-
Strong proficiency with Python, Go, Bash, or a similar language, with demonstrated experience building and maintaining production tooling, automation, CI/CD pipelines, and deployment systems that reduce toil, improve reliability, and simplify ongoing operations.
-
Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and long-term reliability improvements for complex production systems, including the elimination of recurring incidents and operational work.
-
Proven ability to mentor Site Reliability Engineers, help others build stronger technical judgment, communicate clearly with technical and business stakeholders, and lead complex initiatives from planning through delivery.
-
Demonstrated experience applying AI and machine learning to operational data and engineering workflows to identify patterns, forecast reliability or capacity risks, and implement measurable improvements with appropriate safeguards.
You've read the whole posting — now see how you match it.