Crusoe logo

Senior Staff Program Lead, Cloud Engineering Operations

Crusoe

San Francisco, CAFull-time$230–280K/yrPosted todayStill listed today

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$230–280K/yr
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Crusoe is hiring a Senior Staff Program Lead to build and run Engineering Operations for its growing Cloud Engineering organization. The role owns operational programs, metrics, executive reviews, and accountability mechanisms that help teams deliver cloud capacity reliably and meet customer commitments. It begins as an individual-contributor role, with the lead shaping the function’s growth as its scope expands.

Skills & qualifications

RequiredNice to have

Skills

SLO DefinitionIncident ManagementChange ManagementCapacity TrackingOperational Data AnalysisProgram LeadershipCross-Functional CoordinationAccountabilityData AnalysisMetric SelectionOperating Rhythm DesignAmbiguity ManagementLow-Ego CollaborationTooling DevelopmentTechnical Program ManagementSoftware EngineeringSite Reliability EngineeringTechnical Product ManagementAI/ML InfrastructureIncident.ioOpsgeniePagerDuty

Qualifications

10+ Years ExperienceStaff-Level Program LeadershipHigh-Growth Infrastructure ExperienceCloud Provider ExperienceSLO Program ExperienceIncident Review ExperienceChange Management Program Experience

Benefits

Paid Time Off
Medical Insurance
Dental Insurance
Vision Insurance
Parental Leave
Tuition Assistance
401(k) Match

Full job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About This Role Crusoe's Cloud Engineering organization is 380 people today and hiring toward 550. When we sell capacity, we make customers two promises: it will be ready on the date we said, and it will work while they use it. At this scale, keeping those promises doesn't happen by accident anymore. This role exists to make sure it happens on purpose. You'll lead Engineering Operations: the operating system that lets Cloud Engineering leadership agree on outcomes, execute, and keep its promises to customers. You'll own the standing programs that make operations better every week (SLOs, incident follow-up, change safety, capacity delivery tracking, and operational data) and build the systems that feed the weekly executive Engineering Operations review. You don't just report when a metric slips; you find the owner, hold them to the commitment, and make sure a decision happens fast. You'll start as the first hire in this function, working solo while you prove out the model above. As the mandate holds up, you'll shape how the function grows: what to build, who to bring in, and what to keep lean. This isn't a role scoped to a single program or a single dashboard. You will be a strategic advisor to leadership on whether teams are actually executing operationally, and the person who makes clear owners, shared metrics, and a steady cadence stick across the org. The ideal candidate has a technical background (software engineering, SRE, Technical Program Management, or product in an infrastructure context) and is close enough to production systems to hold senior engineering leaders to what they committed.

What You'll Be Working On Engineering Operations programs

  • Own the standing Engineering Operations programs: SLO/SLI definition and attainment, incident management and follow-up, change safety (change policy, maintenance readiness, and pre-production checks), capacity delivery tracking, and operational data. Partner with the engineering leaders who own each program and hold the line on outcomes.

  • Define and track the metrics that show how the org runs, grows, builds, and enables itself: usable capacity, SLO attainment, MTTD and MTTR, customer-found incidents, on-time capacity delivery, milestone slip, change failure rate, follow-up closure rate, and change policy compliance. Data should tell the story before anyone has to ask.

  • Partner with engineering, SRE, data engineering, and product to keep operational data accurate, with one source of truth for each metric.

  • Partner with engineering, TPM, product, SRE, data scientists, customer success, and data center operations to ensure operations run smoothly across all functional areas.

  • Work toward a single pane of glass that makes the operational health of the organization easy to understand at a glance.

Accountability and operating cadence

  • Own the weekly executive Engineering Operations review: build the systems, data, and pre-reads that feed it, and make sure every risk on the page has an owner, a date, and a next step.

  • Run the weekly operations planning forum, and partner with Product on a monthly business review that ties operational health to customer outcomes.

  • Make ownership explicit: clear accountable owners, swimlanes, and articulated outcomes for every program and metric.

  • Hold teams to what they committed. Drive incident follow-ups, overdue actions, and slipping milestones to closure, and escalate fast when an owner can't fix the problem alone.

Growing the function

  • Operate as an individual contributor first: prove out the programs, metrics, and cadence before asking for headcount.

  • Make the case for more capacity as scope outgrows one person, then help hire and onboard the people who join the function.

  • Own the evolution of the operating system itself, not just the programs inside it: retire process that no longer earns its cost, and keep the overhead on engineering teams as low as possible.

What You'll Bring to the Team

  • 10+ years in software engineering, SRE, technical program management, or a technical product role, close enough to production systems to know what an SLO breach actually means.

  • A track record of leading org-wide programs at the staff level or above, and of driving accountability across senior engineering leaders without formal authority.

  • Comfortable starting as a team of one: you don't need a team in place to start driving impact, and you know when the case for headcount is real versus premature.

  • You've operated in high-growth infrastructure environments where processes are still being built; ambiguity doesn't paralyze you.

  • You're a natural coordinator who works across teams without formal authority. Engineering, SRE, and product leads trust you because you follow through.

  • You can turn messy, multi-source data into a clear picture of organizational health, and you know how to pick the few metrics that drive decisions over the many that fill dashboards.

  • You've designed operating rhythms (weekly reviews, business reviews, incident reviews) that leaders actually use, and you keep them light for the teams that feed them.

  • Scrappy, low-ego, high-drive. You build the program and tooling yourself when it doesn't exist yet, and you care more about the outcome than the credit.

Bonus Points

  • Time inside AWS, GCP, Azure, CoreWeave, Lambda Labs, or a similar cloud provider.

  • Experience running an SLO, incident review, or change management program at a cloud or infrastructure company.

  • Familiarity with incident.io, Opsgenie, PagerDuty, or similar incident management platforms at scale.

  • Background in AI/ML infrastructure.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range Compensation will be paid in the range of up to $230,000 - $280,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data. Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.