Vertafore logo

Director, Site Reliability (SRE, SLI/ SLO, Monitoring, Automation)

Vertafore

IN, USAJobNo compensation foundPosted 4mo agoSeen in employer's feed 1 day ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
IN, USA
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Bachelor's degree

Job overview

Vertafore is hiring a Director, Site Reliability (SRE, SLI/ SLO, Monitoring, Automation). The Director of Site Reliability Engineering will lead reliability, performance, and observability initiatives for Vertafore’s product portfolio, owning SLIs/SLOs, incident response, automation, and CI/CD practices while managing multiple teams and collaborating across product development, cloud operations, and security to ensure operational excellence.

Key focus areas include Define and enforce SLIs/SLOs for flagship products, Drive observability strategy across application and infrastructure layers, and Oversee CI/CD pipelines for product deployments using tools such as GitLab and Jenkins.

Successful candidates bring Bachelor's Degree In Computer Science, Bachelor's Degree In Information Systems, and Bachelor's Degree In Related Field. Important skills include SLIs/SLOs, Incident Management, Automation, CI/CD Practices, Observability Strategy, and GitLab. Preferred (not required): Financial Services Background, Healthcare Background, and Regulated Industries Background.

Skills & qualifications

RequiredNice to have

Skills

SLIs/SLOsIncident ManagementAutomationCI/CD PracticesObservability StrategyGitLabJenkinsAnsibleLaunchDarklyInfrastructure as CodeTerraformAWS CloudFormationAWS CDKCapacity PlanningOS PatchingLoad BalancingALBF5CoachingSoftware Engineering PrinciplesAWS KnowledgeContainer OrchestrationLeading Reliability ProgramsArchitecting ApplicationsArchitecting InfrastructureB2B SaaS EnvironmentsLarge-Scale Distributed SystemsCommunicationInfluencingDriving Operational ExcellenceFinancial Services BackgroundHealthcare BackgroundRegulated Industries BackgroundCDKAWSSLI/SLOMonitoringObservabilityCI/CDLeadership

Qualifications

Bachelor's Degree in Computer Science or Information Systems or Related Field18+ Years in Software Engineering, SRE, DevOps, or Reliability Roles5+ Years in Leadership (Director)Background Supporting Financial Services, Healthcare, or Regulated Industries

Full job description

The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role owns SLIs/SLOs, incident response, automation, and CI/CD practices for assigned product families. Directors will manage multiple teams and collaborate with Product Development, Cloud Operations, Information Security, and other SRE leaders to ensure operational excellence.

Key Responsibilities

  • Product Reliability Leadership

  • Define and enforce SLIs/SLOs for a subset of Vertafore flagship products.

  • Drive observability strategy across application and infrastructure layers.

  • Release Engineering & Automation

  • Oversee CI/CD pipelines for product deployments using tools like GitLab, Jenkins, Ansible, LaunchDarkly.

  • Implement Infrastructure-as-Code (Terraform, AWS CloudFormation/CDK) for application provisioning.

  • Incident Management

  • Define 24x7 on-call rotations for assigned products; ensure rapid resolution and blameless postmortems.

  • Cross-Functional Collaboration

  • Partner with Cloud Ops on capacity planning, OS patching (app tier), and load balancing (ALB, F5).

  • Align reliability goals with product roadmaps and customer SLAs.

  • Team Leadership

  • Manage a group of Managers and Engineers; mentor teams on automation, observability, and reliability best practices.

Qualifications

  • Bachelor’s degree in Computer Science, Information Systems, or related field.

  • 18+ years in Software Engineering, SRE, DevOps, or reliability roles; 5+ years in leadership(Director).

  • Proven ability to leverage software engineering principles and practices to solve reliability and operational challenges.

  • Expertise in SLI/SLO and monitoring.

  • Expertise in CI/CD, observability, and incident response.

  • Strong AWS knowledge and experience with container orchestration.

  • Proven ability to lead reliability programs across multiple SaaS products.

  • Experience architecting applications or infrastructure for highgrowth cloud platforms.

  • Experience in B2B SaaS environments involving large-scale distributed systems.

  • Proven leadership communicating and influencing at team, peer, and leadership levels.

  • Demonstrated experience driving operational excellence through metrics and KPIs.

  • (Preferred) Background supporting financial services, healthcare, or regulated industries.

You've read the whole posting — now see how you match it.