Picogrid logo

Site Reliability Engineer

Picogrid

El Segundo, CAFull-time$170–195K/yrPosted 1mo agoVerified open 1w ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$170–195K/yr
Location
El Segundo, CA
Schedule
Full-time
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Security clearance

Job overview

Picogrid is hiring a Site Reliability Engineer. Picogrid is a venture‑backed defense technology company building essential infrastructure to unify sensors, autonomy, and operators for national security, delivering an operational advantage to the United States and its allies. The role focuses on production reliability across cloud and edge, ensuring systems are dependable for warfighters in challenging battlefield conditions.

Key focus areas include Own, define and drive reliability SLIs and SLOs for cloud deployments, Own, define and drive reliability SLIs and SLOs for edge devices deployed in remote and contested areas, and Own the observability stack: Grafana, Prometheus, Loki, and OpenTelemetry with dashboards versioned in git and alerting rules checked in alongside code.

Successful candidates bring 3+ Years SRE Experience, US Citizen Or National, and US Lawful, Permanent Resident. Important skills include Observability, Incident Management, Node Lifecycle, Stateful Workloads, On-Call Culture, and Alerting And Monitoring. Preferred (not required): GovCloud, FIPS, Regulated Or Air-Gapped Environment Experience, and NVIDIA Jetson Platforms.

Skills & qualifications

RequiredNice to have

Skills

ObservabilityIncident ManagementNode LifecycleStateful WorkloadsOn-Call CultureAlerting and MonitoringDevSecOpsSLIs and SLOsGrafanaPrometheusLokiOpenTelemetryGitLog-First TroubleshootingBlameless PostmortemsInfrastructure as CodeKubernetes OperationsWorkload SchedulingStatefulSetsGraceful DrainsLive Cluster DebuggingHigh Signal-to-Noise Ratio Alerting RulesIncident ResponderMethodical Evidence-First TriageTerraformOpenTofuAWSIAMNetworkingMulti-Account EnvironmentsAccount and Workload HardeningHigh Availability Database DeploymentsIoT or Edge Fleet OperationOperating in Scrappy, Fast-Paced EnvironmentsTurning Ambiguous Requirements Into Concrete SolutionsOptimizing for Providing Value Early in ProjectsShort Iteration CyclesGovCloudFIPSRegulated or Air-Gapped Environment ExperienceNVIDIA Jetson PlatformsShared CPU and GPU MemoryThermal ConstraintsNebulaWireGuardTailscaleSLO and Error-Budget ToolingSlothPyrra

Qualifications

3+ Years SRE ExperienceUS Citizen or NationalUS Lawful, Permanent ResidentRefugee Under 8 U.S.C. § 1157Asylee Under 8 U.S.C. § 1158Eligible to Obtain Required Authorizations From US Department of StateActive Security Clearance

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match
Paid Time Off
Parental Leave
Relocation Assistance

Full job description

Who we are

Picogrid is a leading venture-backed defense technology company founded to bridge the decades-long gap between modern technology and the critical demands of national security. Today, we're building the essential infrastructure to unify sensors, autonomy, and operators with our technology deployed in active operations around the world. Our mission is to deliver an operational advantage to secure the United States and its allies.

About the Role As Picogrid's first Site Reliability Engineer you will own production reliability across cloud and edge, from observability and incident response through node lifecycle, stateful workloads, and a fleet of hardware edge devices in the field. You will help build and define the systems, processes and best practices that ensure Picogrid's systems can be relied upon by our warfighters in even the toughest battlefield conditions. You will work with engineers to build a strong on-call culture where issues are root caused swiftly, and ensure our alerting and monitoring have exceptional coverage and signal-to-noise ratio.

Security is a shared responsibility across all our DevSecOps roles, and as part of a scrappy startup team you will be expected to help stand up new infrastructure and other related DevSecOps tasks as needed.

Responsibilities

  • Own, define and drive our reliability SLIs and SLOs for cloud deployments

  • Own, define and drive our reliability SLIs and SLOs for our edge devices deployed in remote and sometimes contested areas

  • Own the observability stack: Grafana, Prometheus, Loki, and OpenTelemetry, with dashboards versioned in git and alerting rules checked in alongside the code they watch

  • Participate in on-call and incident response: log-first troubleshooting, blameless postmortems, and follow-up hardening

  • Encode reliability into infrastructure as code

Required Qualifications

  • 3+ years of experience as an SRE or related roles

  • Deep Kubernetes operations experience: node lifecycle, workload scheduling, StatefulSets, graceful drains, and live cluster debugging

  • Experience designing comprehensive observability dashboards and high signal-to-noise ratio alerting rules

  • You are a competent and experienced incident responder practicing methodical evidence-first triage, blameless postmortems, and turning incidents into durable guardrails

  • Production Terraform or OpenTofu experience

  • Fluent in AWS including IAM, networking, multi-account environments, and account and workload hardening

  • Experience managing high availability database deployments

  • IoT or edge fleet operation experience

  • Comfortable operating in scrappy, fast-paced environments, and turning ambiguous requirements into concrete solutions

  • You optimize for providing value early in projects and short iteration cycles

Preferred Qualifications

  • GovCloud, FIPS, or other regulated or air-gapped environment experience

  • Constrained edge hardware such as NVIDIA Jetson platforms (AGX Thor, Orin Nano), including shared CPU and GPU memory and thermal constraints

  • Overlay or mesh networking operations: Nebula, WireGuard, Tailscale, or similar

  • Standing up SLO and error-budget tooling (sloth, Pyrra, or equivalent) from scratch

  • Active security clearance

Compensation & Benefits

  • Base salary range: $170,000 - $195,000 per year. Base salary is just one part of your total compensation package at Picogrid.

  • Significant stock options with a high potential upside as an early-stage company

  • 401(k) with employer matching

  • Full health coverage (medical, dental, and vision insurance)

  • Relocation assistance provided (if applicable)

  • Unlimited PTO (two-week minimum) and 11 paid holidays per year

  • Paid parental leave for both parents

  • Lunch provided when working in-office and a fully stocked kitchenette

  • Free EV charging at the HQ

  • Unique office in El Segundo, CA stocked with quality coffee, snacks, and craft beer

Export Control Requirements To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.

Equal Employment Opportunity (EEO) Policy Picogrid is committed to providing a professional work environment free from discrimination, harassment, and retaliation. We are an equal opportunity employer and make all employment decisions based on merit, qualifications, and business needs.

#LI-DNP

Equal Employment Opportunity (EEO) Policy

Picogrid is committed to providing a professional work environment free from discrimination, harassment, and retaliation. We are an equal opportunity employer and make all employment decisions based on merit, qualifications, and business needs.

You've read the whole posting — now see how you match it.