HPC Infrastructure and Cluster Engineer
Arena Technical Resources, LLC
Springfield, VAJob$180–200K/yrSeen 1mo agoSeen in employer's feed 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Springfield, VA, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
Arena Technical Resources seeks an Infrastructure & Cluster Engineer to manage administration, health, and performance of a dedicated customer compute cluster, ensuring high availability, security, and optimized hardware for AI/ML workloads.
Skills & qualifications
Skills
Qualifications
Full job description
HPC Infrastructure and Cluster Engineer
Location: Springfield, VA, US
Job ID: ATR 18077
Job Description
Job Title: HPC Infrastructure and Cluster Engineer
Job Location: Springfield, VA
Compensation: $180,000 - $200,000
Eligibility/Clearance: Candidate must possess an active TS/SCI Clearance
and the ability to obtain a CI Polygraph
Job Description:
We are seeking an Infrastructure & Cluster Engineer to manage the
administration, health, and performance of the foundational compute
environment under the User Facing and Data Center Services (UDS)
contract. In this role, you will be responsible for the end-to-end
administration of a dedicated customer compute cluster. Your primary
mission is to ensure a highly available, secure, and optimized hardware
foundation. By maintaining a robust infrastructure, you will directly
contribute to the critical technology integration and performance
engineering efforts, ensuring a highly reliable platform for integrating
and executing complex customer workloads.
Key Responsibilities
- Cluster Administration: Manage the day-to-day operations of the
customer compute cluster, including Linux operating system
administration, hardware monitoring, patching, and system upgrades.
- Resource and Job Management: Configure, maintain, and optimize
workload management and orchestration platforms, utilizing the
Run:AI job scheduler to ensure efficient distribution of intensive
AI/ML workloads across the cluster.
- Infrastructure Optimization: Tune cluster performance at the
hardware, operating system, and network levels to maximize compute
efficiency and data throughput for customer workloads.
- Storage and Network Management: Administer storage solutions and
high-speed networking fabrics. Support the transition to and ongoing
management of an InfiniBand GPU-to-GPU network infrastructure to
minimize latency for distributed operations.
- Environment Configuration: Partner with technology integration teams
to provision specific environments, dependencies, and container
platforms, specifically leveraging Red Hat OpenShift, required for
seamless customer model deployment.
- Security and Compliance: Ensure all infrastructure components remain
compliant with federal security standards, implementing strict
access controls and maintaining system accreditations.
Skills/Qualifications:
Required:
Experience:
-
Clearance: Active TS/SCI with the ability to obtain CI Poly.
-
Experience: 5+ years of experience in Linux systems administration
and infrastructure management with a specific focus on
high-performance computing environments.
-
Technical Skills:
-
Expertise in managing bare-metal servers, enterprise storage
arrays, and advanced network configurations (Experience with
InfiniBand).
- Strong proficiency with workload managers, job schedulers, and
AI orchestration tools (e.g., Run:AI, SLURM).
- Hands-on experience with enterprise container orchestration
platforms, specifically OpenShift or Kubernetes.
- Experience writing automation and configuration scripts (e.g.,
Bash, Python) to streamline cluster maintenance.
- Troubleshooting Focus: Proven ability to diagnose and resolve
complex hardware, network, and OS-level issues.
Desired:
- Familiarity with parallel file systems and high-throughput storage
architectures.
- Prior experience engineering or managing high-speed GPU-to-GPU
communication topologies.
Education:
Bachelors Degree in Computer Science or a related field
ATR is an Equal Opportunity Employer (EOE) who will provide equal
employment opportunity to employees and applicants for employment
without regard to race, ethnicity, religion, color, sex, pregnancy,
national origin, age, veteran status, ancestry, sexual orientation,
gender identity or expression, marital status, family structure, genetic
information, or mental or physical disability
First Name
Required
Last Name
Required
Email Address
Required
Phone Number
CountryNoneAfghanistanÅland IslandsAlbaniaAlgeriaAmerican SamoaAndorraAngolaAnguillaAntarcticaAntigua and BarbudaArgentinaArmeniaArubaAustraliaAustriaAzerbaijanBahamasBahrainBangladeshBarbadosBelarusBelgiumBelizeBeninBermudaBhutanBoliviaBonaire, Sint Eustatius and SabaBosnia and HerzegovinaBotswanaBouvet IslandBrazilBritish Indian Ocean TerritoryBritish Virgin IslandsBruneiBulgariaBurkina FasoBurundiCabo VerdeCambodiaCameroonCanadaCayman IslandsCentral African RepublicChadChileChinaChristmas IslandCocos (Keeling) IslandsColombiaComorosCongoCongo-BrazzavilleCook IslandsCosta RicaCôte d'IvoireCroatiaCubaCuraçaoCyprusCzechiaDemocratic People's Republic of KoreaDenmarkDjiboutiDominicaDominican RepublicEcuadorEgyptEl SalvadorEquatorial GuineaEritreaEstoniaEthiopiaFalkland IslandsFaroe IslandsFederated States of MicronesiaFijiFinlandFranceFrench GuianaFrench PolynesiaFrench Southern TerritoriesGabonGambiaGeorgiaGermanyGhanaGibraltarGreeceGreenlandGrenadaGuadeloupeGuamGuatemalaGuernseyGuineaGuinea-BissauGuyanaHaitiHeard Island and McDonald IslandsHondurasHong KongHungaryIcelandIndiaIndonesiaIraqIrelandIslamic Republic of IranIsle of ManIsraelItalyJamaicaJapanJerseyJordanKazakhstanKenyaKiribatiKuwaitKyrgyzstanLao People's Democratic RepublicLatviaLebanonLesothoLiberiaLibyaLiechtensteinLithuaniaLuxembourgMacaoMacedoniaMadagascarMalawiMalaysiaMaldivesMaliMaltaMarshall IslandsMartiniqueMauritaniaMauritiusMayotteMexicoMonacoMongoliaMontenegroMontserratMoroccoMozambiqueMyanmarNamibiaNauruNepalNetherlandsNew CaledoniaNew ZealandNicaraguaNigerNigeriaNiueNorfolk IslandNorthern Mariana IslandsNorwayOmanPakistanPalauPanamaPapua New GuineaParaguayPeruPhilippinesPitcairnPolandPortugalPuerto RicoQatarRepublic of KoreaRepublic of MoldovaReunionRomaniaRussiaRwandaSaint BarthelemySaint Helena, Ascension and Tristan da CunhaSaint Kitts and NevisSaint LuciaSaint MartinSaint Pierre and MiquelonSaint Vincent and the GrenadinesSamoaSan MarinoSao Tome and PrincipeSaudi ArabiaSenegalSerbiaSeychellesSierra LeoneSingaporeSint Maarten (Dutch part)SlovakiaSloveniaSolomon IslandsSomaliaSouth AfricaSouth Georgia and the South Sandwich IslandsSouth SudanSpainSri LankaState of PalestineSudanSurinameSvalbard and Jan MayenSwazilandSwedenSwitzerlandSyriaTaiwanTajikistanThailandTimor-LesteTogoTokelauTongaTrinidad and TobagoTunisiaTurkeyTurkmenistanTurks and Caicos IslandsTuvaluU.S. Virgin IslandsUgandaUkraineUnited Arab EmiratesUnited KingdomUnited Republic of TanzaniaUnited StatesUnited States Minor Outlying IslandsUruguayUzbekistanVanuatuVaticanVenezuelaVietnamWallis and FutunaWestern SaharaYemenZambiaZimbabwe
State/ProvinceNoneAlabamaAlaskaArizonaArkansasCaliforniaColoradoConnecticutDelawareFloridaGeorgiaHawaiiIdahoIllinoisIndianaIowaKansasKentuckyLouisianaMaineMarylandMassachusettsMichiganMinnesotaMississippiMissouriMontanaNebraskaNevadaNew HampshireNew JerseyNew MexicoNew YorkNorth CarolinaNorth DakotaOhioOklahomaOregonPennsylvaniaRhode IslandSouth CarolinaSouth DakotaTennesseeTexasUtahVermontVirginiaWashingtonWashington, D.C.West VirginiaWisconsinWyoming
City
ZIP/Postal Code
Resume
Choose File...
Required, maximum file size is 512KB, allowed file types are doc, docx, pdf, odf, and txt
Message
Success!
Your application was successfully sent!
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Data Center Infrastructure EngineerPeraton · Washington, DC · $135–216K/yrPosted 2w agoPosted 2w ago
Infrastructure LeadSteampunk · McLean, VA (Hybrid) · $140–180K/yrPosted 5 days agoPosted 5 days ago
Senior Strategy Advisor (Digital Infrastructure)Development Finance Corporation · Washington, DC · $144–187K/yrPosted 1w agoPosted 1w ago
Senior Solution Architect, AI InfrastructureNVIDIA · Remote · US · $184–288K/yrPosted 4w agoPosted 4w ago
Senior Strategy Advisor (Transport Infrastructure)Development Finance Corporation · Washington, DC · $144–187K/yrPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.