Stanford University logo

Research Computing Systems Engineer

Stanford University

Stanford, CAJobSeen todaySeen in employer's feed today

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Stanford, CA
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Stanford Research Computing seeks a systems engineer to steward shared infrastructure supporting large-scale HPC clusters and petabyte-scale storage used by researchers across Stanford. The role leads backbone services including bastion hosts, network license services, observability, container hosts, and cluster service nodes. It also contributes to a co-location initiative, strengthens resilience and security, and owns hardware lifecycle planning and vendor engagement.

Skills & qualifications

RequiredNice to have

Skills

Bastion Host AdministrationSystems HardeningHigh Availability DesignFLEXlmFirewall PolicyPrometheusGrafanaSplunkXDMoDNetwork ArchitectureContainer HostsAnsibleGitConfiguration ManagementSecure Build StandardsEndpoint Agent DeploymentLogging Agent DeploymentCredential ManagementAccess Lifecycle ManagementHardware Lifecycle PlanningCapital PlanningVendor Management

Full job description

Job Description

Stanford Research Computing is looking for a talented systems engineer to join our team of collaborative and innovative professionals helping Stanford's faculty and students use advanced computing and data tools to explore new frontiers in knowledge and solve some of humanity's most urgent problems. Our staff work directly with some of the world's top researchers in a broad range of disciplines, across all of Stanford's seven schools — while also supporting and learning from each other in cross-project endeavors. We maintain and steadily improve an advanced research computing facility, and we support a variety of environments for Stanford research. In Stanford Research Computing, you'll have a rare opportunity to contribute to discoveries and inventions that have global reach and positive impact, and to share in the curiosity and commitment of the scholars and scientists who lead these projects.

About the Role

Stanford Research Computing operates a portfolio of large-scale HPC clusters and petabyte-scale storage systems serving thousands of researchers. Every one of those platforms depends on a layer of shared infrastructure that researchers never see: the bastion hosts our administrators pass through to reach management networks, the network license servers that let research software start, the container hosts carrying our internal services, and the observability stack that tells us something is wrong before a researcher has to.

In this role you will be the primary steward of that layer, serving as lead for the "backbone" services the rest of our environment depends on. You will also contribute to a co-location initiative, standing up management infrastructure and geographically distinct replication in a remote data center. The work spans a wide range of technologies, with substantial engineering ahead in hardware lifecycle, redundancy, and configuration management.

Why Stanford? You won't be maintaining back-office IT. You will be building the connective tissue that a multi-petabyte, multi-platform research computing environment runs on — the systems that decide whether a downtime is a scheduled inconvenience or a lost quarter of someone's research.

Responsibilities

  • Bastion hosts: Lead the administration, hardening, and modernization of bastion and jump hosts across multiple data centers, and design the next generation with redundancy and high availability. Support partner unit administrators moving from VDI to bastion-based access.

  • Network license services: Operate and improve the primary and secondary FLEXlm license servers behind commercial research and statistical software. Right-size firewall policy, retire orphaned license hosts, and bring runbooks and user documentation current.

  • Central observability service: Lead the central observability service and the servers behind it — Prometheus, Grafana, Splunk, XDMoD, and facility telemetry. Consolidate legacy monitoring and build alerting that reaches the right person with enough context to act.

  • Co-location project: Support a funded co-location initiative from design through production hand-off: network architecture, management infrastructure and a bastion host in the remote facility, and a geographically distinct replication target for independent backups.

  • Container hosts and backbone services: Lead the container hosts and the internal services they carry, along with cluster head and service nodes, bare-metal provisioning, and the VMs behind other Research Computing backbone services. Bring these under Ansible and Git-based configuration management.

  • Security and compliance: Establish and maintain secure build standards across the systems you own, in line with university, regulatory, and contractual requirements. Deploy endpoint and logging agents, remediate scan findings, and manage credential and access lifecycle.

  • Resilience and documentation: Reduce single points of failure, and document core service requirements and details so that any member of the team can operate, restore, or rebuild a service.

  • Planning and hardware lifecycle: Meet regularly with platform leads and service owners to understand upcoming needs, own hardware lifecycle and capital planning for the infrastructure you lead, and help develop business cases for replacement and expansion.

  • Vendor engagement : Liaise with hardware and software vendors and support partners to triage and resolve issues, manage RMAs, and keep systems under appropriate warranty and support coverage.

At Stanford, every employee plays a part in our mission for a better tomorrow—from researchers to operations and from food services to educators. Join our team to work towards a future you believe in.

The job duties listed are typical examples of work performed by positions in this job classifications and are not designed to contain or be interpreted as a comprehensive inventory of all duties, tasks and responsibilities. Specific duties and responsibilities may vary depending on department or program needs without changing the general nature and scope of the job or level of responsibility. Employees may also perform other duties as assigned.

Consistent with its obligations under the law, the University will provide reasonable accommodations to applicants and employees with disabilities. Applicants requiring a reasonable accommodation for any part of the application or hiring process should contact Stanford University Human Resources by submitting a help ticket (https://stanford.service-now.com/humanresources\_services?id=sc\_cat\_item&sys\_id=aa4161da130d574019813598d144b0b2) .

Stanford is an equal employment opportunity and affirmative action employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by law. See Admin Guide 1.7.4: Equal Employment Opportunity, Non-Discrimination, and Affirmative Action Policy; Know Your Rights: Workplace Discrimination is Illegal .

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.