Sieve logo

Member of Technical Staff, Reliability

Sieve

San Francisco, CAFull-time$150–300K/yrPosted 7mo agoStill listed 3w ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
$150–300K/yr
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Sieve is hiring a Member of Technical Staff, Reliability. Sieve is a multi‑modal lab building the world’s highest‑quality training datasets across video, audio, images, text, and 3D, processing petabytes of video at scale. The company seeks a highly‑ownership engineer to design, validate, and harden infrastructure for PB‑scale workloads, own incident response, and develop monitoring, security, and CI/CD tooling for the entire engineering team.

Key focus areas include Design and validate infrastructure for petabyte‑scale workloads, Own incident response for Sev 0/Sev 1 production incidents, and Harden systems against failure and improve reliability.

Successful candidates bring 3+ Years Building Internal Infrastructure At Scale, Experience On-Call For Sev 0 / Sev 1 Production Incidents, and Onsite In San Francisco 5 Days Per Week. Important skills include GCP, AWS, Oracle, Cloudflare, Infrastructure as Code, and Python. Preferred (not required): Reliability, Incident Management, Monitoring, and CI/CD Tooling.

Skills & qualifications

RequiredNice to have

Skills

GCPAWSOracleCloudflareInfrastructure as CodePythonGoRustC++Networking IntuitionSecurity IntuitionSSO ImplementationReliabilityObservabilityIncident ManagementMonitoringCI/CD ToolingThroughputAnticipate Failure ModesEliminate Operational RiskDesign Systems That Don't BreakTerraformArgoHelmKustomizePrometheusOpenTelemetryVictoriaMetricsBuilding Lightweight Internal ToolingAPIDashboardsSvelteObject Storage SystemsActive GitHubPortfolio Projects

Qualifications

3+ Years Building Internal Infrastructure at ScaleExperience on-Call for Sev 0 / Sev 1 Production IncidentsOnsite in San Francisco 5 Days Per WeekL3 on-Call Experience

Benefits

Medical Insurance
401(k) Match
Dental Insurance
Vision Insurance

Full job description

ABOUT US

Sieve is a multi-modal lab curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data.

We partner with top AI labs and did $XXM last quarter alone, as a team of ~30 people. We also raised our Series A from Tier 1 firms such as Matrix Partners https://matrix.vc/, Swift Ventures https://www.swift.vc/, Y Combinator https://www.ycombinator.com/, and AI Grant https://aigrant.com/.

 

WHY NOW

Sieve is one of the most capital-efficient teams in AI — roughly 30 people serving the world's leading AI labs across every major data modality. You'll join early, own problems end-to-end, and watch your work ship directly into the models defining the frontier.

ABOUT THE ROLE

We process petabytes of video across thousands of nodes and multiple cloud environments, and as we scale, reliability, observability, and security become existential. We're hiring our first engineer fully dedicated to Sieve's infrastructure foundation — a high-ownership role working directly with our CTO and founding engineers to build the core tooling that powers all of engineering. You'll design and validate the infrastructure behind PB-scale workloads, own incident response, harden systems against failure, and build the monitoring, security, and CI/CD tooling the whole team relies on.

This role is ideal for someone who thinks deeply about reliability, throughput, observability, and security — the kind of engineer who anticipates failure modes, eliminates operational risk, and designs systems that don't break. If something goes down, you take it personally, and you thrive in that level of responsibility.

REQUIREMENTS

  • 3+ years building internal infrastructure at scale

  • Experience on-call for Sev 0 / Sev 1 production incidents (L3 preferred)

  • Strong cloud experience (GCP, AWS, Oracle, Cloudflare, etc.)

  • Deep Infrastructure-as-Code experience (Terraform preferred)

  • Familiarity with Argo, Helm, Kustomize, or similar deployment tools

  • Experience operating observability systems (Prometheus, OTel, VictoriaMetrics)

  • Backend fundamentals in Python, Go, Rust, or C++

  • Strong networking + security intuition, including SSO implementation

  • In-person at our SF HQ

  • Bonus: Experience building lightweight internal tooling (APIs, dashboards, Svelte)

  • Bonus: Familiarity with object storage systems ("buckets")

  • Bonus: Active GitHub or portfolio projects

BENEFITS

  • 401k + Full Health Insurance

  • Breakfast, Lunch, and Dinner covered and your choice of snacks

  • Ubers covered home

*all roles at Sieve require you to be onsite in San Francisco 5 days per week

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.