University of North Carolina- Chapel Hill logo

HPC DevOps Engineer

University of North Carolina- Chapel Hill

Bolton, MSHybridJob$105–110K/yrTracked 1mo agoSeen in employer's feed 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$105–110K/yr
Location
Bolton, MSHybrid
Work Authorization
Not specified

Job overview

University of North Carolina- Chapel Hill is hiring a HPC DevOps Engineer. The HPC DevOps Engineer will design, deploy, operate, and maintain services and solutions for leading-edge academic research, primarily related to High Performance Computing (HPC) and High Throughput Computing (HTC). This role involves contributing to the design, build, and operation of HTC/HPC services and ancillary systems. The ideal candidate will have a comprehensive understanding of Linux systems administration in HPC environments and strong technical capabilities.

Key focus areas include Design, deploy, operate, and maintain services and solutions for academic research, Lead and contribute to the design, build, and operation of HTC/HPC services, and Manage high-speed storage and networking, containers, databases, monitoring, and orchestration services.

Successful candidates bring Experience Managing InfiniBand HPC Clusters, Experience Managing Multi Petabyte Enterprise Storage Platforms, and Hands On Data Center Hardware Ecosystems Experience. Important skills include HPC, HTC, High-Speed Storage, Networking, Containers, and Databases.

Skills & qualifications

RequiredNice to have

Skills

HPCHTCHigh-Speed StorageNetworkingContainersDatabasesMonitoringOrchestration ServicesLinux Systems AdministrationLeadershipCommunicationInfiniBandCobblerSaltStackAnsibleRHELPrometheusGrafanaElastic StackOpenSearchSplunkAlertmanagerLinux Operating SystemsAutomationConfiguration ManagementEnterprise Storage PlatformsLogging Data InterpretationAlerting Data InterpretationDiagnosticProblem-SolvingChange ManagementOperational Best PracticesVendor CoordinationService Provider CoordinationInternal Stakeholder CoordinationTechnical Information ConveyanceTeam CollaborationControlled Information HandlingRegulated Data HandlingSecurityAccess ControlLarge-Scale OS DeploymentsIncident ManagementSystem Behavior AnalysisOrganizational AwarenessOperational Judgment

Qualifications

Experience Managing InfiniBand HPC ClustersExperience Managing Multi Petabyte Enterprise Storage PlatformsHands on Data Center Hardware Ecosystems ExperienceFamiliarity With Infrastructure AutomationExperience Maintaining RHEL Based Environments at ScaleBackground Working With Observability PlatformsAbility to Lift and Maneuver Equipment Up to 50 PoundsAbility to Work in Hot/Cold Aisle ConditionsAbility to Perform Tasks in Confined Rack EnvironmentsCapacity to Stand, Bend, and Work in Data-Center Spaces for Extended Periods

Full job description

Employment Type: Permanent Staff (EHRA NF)

Vacancy ID: NF0009910

Salary Range: $105,000 - $110,371

Position Summary/Description:

This position may be eligible for a hybrid work arrangement that may include a partially remote work location, consistent with System Office policy. UNC Chapel Hill employees are generally required to reside within a reasonable commuting distance of their assigned duty station.

The HPC DevOps Engineer will have a broad role within Research Computing at UNC Chapel Hill, designing, deploying, operating and maintaining services and solutions in support of leading -edge academic research, primarily related to High Performance Computing ( HPC ) as well as High Throughput Computing ( HTC ).

Responsibilities of the HPC DevOps Engineer include leading and contributing to a variety of areas within Research Computing related to the design, build, and operation of HTC / HPC services along with ancillary systems including high-speed storage and networking, containers, databases, monitoring, orchestration services, etc.

The ideal candidate for this position should possess a comprehensive understanding of Linux systems administration in HPC environments, broad technical capabilities, and enjoy applying these skills in academic research to further the mission of the University.

The successful candidate should possess excellent leadership and communication skills to effectively lead complex projects. They should have the ability to engage directly with faculty and researchers, as well as leverage research computing communities of practice beyond Carolina.

ITS Research Computing aims to provide world-class computing and data infrastructure as well as other services and capabilities in support of research needs for faculty, staff, students, and collaborators. Our goal is to provide state -of-the-art environments and services supporting the highest level of multidisciplinary research.

Education and Experience:

Experience managing InfiniBand HPC clusters. Experience managing multi petabyte enterprise storage platforms, including vendor specific administration tools, performance tuning concepts, and lifecycle planning. Hands on experience with data center hardware ecosystems such as multi node server deployments, high density rack environments, and hardware lifecycle management. Familiarity with infrastructure automation, including developing or extending provisioning and configuration management code in tools such as Cobbler, SALT , Ansible, or similar platforms. Experience maintaining RHEL based environments at scale, including custom repository management, patch orchestration, and compliance aligned OS baselines. Background working with observability platforms (e.g., Prometheus, Grafana, ELK /Opensearch, Splunk, Alertmanager)

and contributing to monitoring or logging architecture.

Essential Skills:

  • Proficiency with Linux operating systems, including installation, configuration, troubleshooting, and lifecycle management in an enterprise or research-computing context. Experience with automation and configuration-management tools such as Cobbler, SALT , or comparable platforms used to deploy and manage systems at scale. Familiarity with enterprise storage platforms, including routine configuration tasks, capacity monitoring, health assessment, and basic hardware maintenance such as drive replacement. Ability to interpret and act on monitoring, logging, and alerting data to maintain operational continuity across diverse service lines. Understanding of RHEL repo management and practices for maintaining consistent, secure OS-layer operations.

  • Ability to lift and maneuver equipment up to 50 pounds with or without reasonable accommodations, work in hot/cold aisle conditions, and perform tasks in confined rack environments. Capacity to stand, bend, and work in data-center spaces for extended periods while performing installation, cabling, and hardware maintenance tasks.

  • Strong diagnostic and problem-solving abilities, including the capacity to assess hardware and system issues under time constraints. Demonstrated adherence to change-management and operational best practices in complex technical environments. Ability to coordinate effectively with vendors, service providers, and internal stakeholders to support hardware and storage operations.

  • Clear written and verbal communication skills, including the ability to convey technical information to both technical and non-technical audiences. Ability to work collaboratively within a team where some responsibilities are shared and others are independently owned.

  • Understanding of operational practices required to support environments involving controlled information and regulated data, including attention to security, access control, and data-handling expectations.

  • The position requires expertise with Linux operating systems and the ability to engineer, maintain, and troubleshoot large-scale OS deployments for HPC and research-computing environments. The role demands sustained focus, careful change-management discipline, and the ability to diagnose complex system-level issues in a fast-moving research environment.

  • Duties include contributing to the architecture and automation that underpin observability platforms, ensuring reliable telemetry collection and correlation across diverse service lines, and interpreting operational signals to maintain service health, security posture, and compliance expectations. The role requires disciplined incident response, careful analysis of system behavior, and the ability to translate monitoring insights into stable, well-governed operations.

  • The role requires strong organizational awareness, clear communication with service owners and vendors, and steady operational judgment to maintain reliable, well-governed storage services.

AA/EEO Statement:

The University is an equal opportunity employer and welcomes all to apply without regard to age, color, gender, gender expression, gender identity, genetic information, national origin, race, religion, sex, or sexual orientation. We encourage all qualified applicants to apply, including protected veterans and individuals with disabilities.

You've read the whole posting — now see how you match it.