Filevine logo

Staff Site Reliability Engineer ›

Filevine

RemoteRemoteJobNo compensation foundPosted 3w agoVerified open 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
RemoteRemote
Work Authorization
Not specified

Job overview

Filevine is hiring a Staff Site Reliability Engineer ›. Filevine seeks a Staff Site Reliability Engineer who will act as the senior technical authority on the SRE team, shaping engineering culture, defining production standards, and linking business goals with reliable, internet‑scale execution while integrating AI/ML into observability practices.

Key focus areas include Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence, Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems, and Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.

Important skills include Site Reliability Engineering, Cloud Infrastructure, Reliability, Incident Response, Capacity Planning, and Automation. Preferred (not required): FedRAMP, CJIS, HIPAA, and SOC 2.

Skills & qualifications

RequiredNice to have

Skills

Site Reliability EngineeringCloud InfrastructureReliabilityIncident ResponseCapacity PlanningAutomationKubernetesDatadogMentorshipDistributed SystemsObservabilityReliability EngineeringMentoring EngineersInfluence Technical DirectionCommunicate Production RiskAIOpsArtificial IntelligenceMachine LearningAnomaly DetectionIncident ManagementAutomated RemediationResource OptimizationProduction-Scale Problem SolvingSoftwareInfrastructure as CodePlatform CapabilitiesOperational ExcellenceSLIsSLOsError BudgetsOperational ReadinessNew RelicPythonGoBashBuilding Production ToolingFedRAMPCJISHIPAASOC 2PCI DSS

Qualifications

12+ Years Software Engineering Experience6+ Years SRE Experience3+ Years Leading Complex Cross-Functional Technical Initiatives

Full job description

Role Summary As a Staff Site Reliability Engineer at Filevine, you are the senior technical authority on the SRE team and a strategic partner to engineering leadership. You don’t just maintain systems — you shape engineering culture, define the technical standard for how Filevine runs in production, and bridge the gap between high-level business goals and robust, internet-scale technical execution. You bring a forward-looking perspective — actively shaping how AI and machine learning drive the future of reliability practice. You own the roadmap across two critical SRE domains — Observability & Alerting and Platform Infrastructure — and are accountable for ensuring the team solves reliability problems permanently rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering leadership on significant technical decisions, mentor engineers across experience levels, and influence reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the senior technical voice responsible for ensuring that uptime, incident response, and every production change meet the operational standard the business demands. This role does not participate in on-call rotation, but you are deeply invested in the engineers who do — shaping the on-call strategy, tooling, and culture that make production support sustainable and effective. Who You Are The Technical Authority

  • Master of the Craft: You bring deep expertise in distributed systems, cloud infrastructure, observability, and reliability engineering. You raise the technical standard for every engineer around you and thrive where the challenges are complex and the stakes are real.

  • Technical Leader and Mentor: You are passionate about mentoring engineers and investing in their growth. You influence technical direction and communicate production risk clearly across engineering, product, and executive audiences.

  • Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of AI and machine learning in observability, anomaly detection, incident response, automated remediation, and resource optimization.

  • Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into durable solutions for systems where availability, performance, and production changes carry meaningful business impact.

  • Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform capabilities to eliminate toil and make systems safer, more scalable, and easier to operate. What you will do Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.

  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.

  • Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.

  • Lead the organization through complex production incidents and turn post-incident learning into permanent engineering improvements.

  • Build self-service platform capabilities that reduce toil, improve engineering safety and velocity, and make every team more capable of owning their own reliability.

  • Mentor engineers and serve as a trusted technical authority for long-term reliability and platform direction. Qualifications 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE, including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives for distributed production systems.

  • Expert-level depth in observability and platform infrastructure, with broad expertise in incident response, capacity planning, automation, and reliability engineering.

  • Advanced experience with a major container-orchestration platform, preferably Kubernetes, and an observability platform such as New Relic, Datadog, or equivalent.

  • Strong software-engineering ability in Python, Go, Bash, or another general-purpose language, with experience building production tooling, automation, or platform capabilities.

  • Proven ability to mentor engineers and communicate technical risk clearly to engineering, product, and executive audiences.

  • Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.

You've read the whole posting — now see how you match it.