Metriport logo

Staff Data Engineer

Metriport

San Francisco, CAHybridFull-time$200–260K/yrPosted 3w agoVerified open 6 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$200–260K/yr
Location
San Francisco, CAHybrid
Schedule
Full-time
Work Authorization
Not specified

Job overview

Metriport is hiring a Staff Data Engineer. Metriport, a technology company focused on healthcare data exchange, seeks a data engineering leader to own and evolve its end‑to‑end data platform. The role involves shaping architecture, guiding engineers, and delivering reliable, scalable solutions that power clinical data pipelines for millions of patients.

Key focus areas include Set technical direction for data platform and evolve architecture, Drive critical data projects end‑to‑end and ship v0 to v1, and Support AI/ML engineers by providing required data.

Successful candidates bring 8+ Years Of Engineering Experience, Building, Operating, And Scaling Data Platforms, and Processing Terabytes Of Data. Important skills include Distributed Processing, Lakehouses, Warehouses, Streaming, Orchestration, and Technical Guidance. Preferred (not required): HL7.

Skills & qualifications

RequiredNice to have

Skills

Distributed ProcessingLakehousesWarehousesStreamingOrchestrationTechnical GuidanceTechnical DirectionData ReliabilityData QualityData GovernanceCustomer Value DeliveryHacker MindsetETL/ELT ArchitectureBatch ProcessingQuery EnginesDesign DocumentsMentoring EngineersCode ReviewsPRsHands-on CodingSprint PlanningRetrosEngineering Roadmap ContributionOn-Call RotationData Architecture DesignIngestionStorageWarehousingServingCost OptimizationLatency OptimizationCorrectnessOperabilityCloud-Native Data StacksApache SparkParquetIcebergDeltaAmazon S3SnowflakeBigQueryRedshiftDbtApache AirflowDagsterKafkaAWS KinesisSoftware Engineering FundamentalsTypeScriptPythonNode.jsAWSAmazon ECSAWS LambdaAWS SQSAWS SNSBatchCDKPostgreSQLAuroraDynamoDBAthenaSageMakerHL7Data LakeData WarehouseLeadershipCoachingEntrepreneurial MindsetOwnership

Qualifications

8+ Years Engineering ExperienceLocated in San Francisco Bay Area or Willing to RelocateExperience Leading EngineersExperience Building or Supporting ML/Data Science WorkflowsHealthcare Standards Technologies (FHIR, HIE, IHE, EHR/EMR, NPI, TEFCA, ADT, HL7, HEDIS, RAF, SNOMED, LOINC, ICD-10)

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
Paid Time Off
401(k) Match

Full job description

ABOUT US

Medical data exchange is one of the biggest unsolved problems in US healthcare. If you've spent any real time in the healthcare system, you've felt it, and it only gets worse the sicker and older you get. Metriport exists to fix that. We connect to the data sources the healthcare providers of tomorrow need, take raw data that's unusable in its source form, and turn it into a single clean format that care teams, and their agents, actually use to improve outcomes. Record retrieval across thousands of legacy systems and antiquated data pipes that used to take weeks now takes seconds. Clinicians walk into appointments with the full patient picture already in hand, and that speed can mean spotting a condition early enough to actually treat it.

We're not a healthcare company. We're a technology company that happens to operate in healthcare, and we build every layer ourselves: acquisition, transform, and insight. Legacy EHRs were built to get providers paid, not to serve the clinical experience. We're building what should have existed instead: the system of record for all of US healthcare. We started in data exchange. Today we compete with the largest data platforms in the space. Tomorrow we will be the infrastructure layer healthcare runs on.

We've found product-market fit (multi-million dollar ARR, 100+ customers including Amazon One Medical, Circle Medical, Color Health, and Strive Health) and we're backed by top-tier VCs with years of runway ahead of us. If you’ve had a good healthcare experience in the past few years, there’s a good chance Metriport was behind it. That’s the version of healthcare we want people to not just expect, but demand. Join us and build the future.

About You

We're a small, high-output team, mostly former founders including YC alumni, and we operate with real autonomy and almost no bureaucracy. We hire on competence, not pedigree. Leaders are in the office six days a week, and we generally expect the team to be available six days a week too. Not because we count hours, but because there's always more wood to chop when revolutionizing how tomorrow’s healthcare providers deliver care to their patients today. We trust our team to take the time off they need and we've never said no to a time off request.

You have an entrepreneurial mindset and a strong sense of ownership. You don't wait to be told what's broken. Instead, you believe you can figure out any problem that lands in your lap even outside your exact domain. When someone scopes something for three weeks, you ask why it can't be done in three days, and you also know the difference between shipping a fast v0 and cutting a corner that comes back to bite you. You walk the tightrope between craft and speed without falling off either side.

THE ROLE

We're looking for a data engineering leader who can own our data platform end-to-end:

  • You've designed and operated data platforms at scale, with broad hands-on experience across the ecosystem — distributed processing, lakehouses, warehouses, streaming, orchestration — and people usually come to you for technical guidance.

  • You've led engineers before — as a team lead, tech lead, or founder — and you own outcomes: technical direction, results, and the work of the people you lead. You're energized by multiplying a team's output, not just your own.

  • You're entrepreneurial-minded with an olympian-level work ethic (about half our engineering team are former founders).

  • You own data reliability, quality, and governance as first-class parts of delivery.

  • You care about delivering value to customers, not about what frilly new tech is under the hood.

  • When someone scopes a project for 3 weeks, you ask "why can't it be done in 3 days?" — and you help others develop that same instinct.

  • You're a hacker at heart, with a good sense of which rules should, and shouldn't, be broken.

WHAT YOU'LL BE DOING

We ingest clinical data for millions of patients from external healthcare sources, with continuous updates for a growing subset of those patients. You'll own the architecture and evolution of the data platform that powers our product — and ship it to customers fast.

Day to day, that looks like:

  • Setting the technical direction for our data platform: evolving our warehouse, data lake, and ETL/ELT architecture to scale with patient and customer growth, and picking the right tools (batch and streaming processing, table formats, orchestration, query engines).

  • Driving the critical data projects end-to-end: writing Design Documents, shipping v0's quickly, and iterating to v1 and beyond.

  • Supporting AI/ML efforts: making sure the AI Engineers have the data they need.

  • Multiplying the team: mentoring engineers on data fundamentals, reviewing designs and PRs, and judging when to invest in quality vs. ship fast.

  • Eventually, acting as Team Lead for a group of engineers: breaking down and delegating work, unblocking teammates, and owning your team's delivery — while staying hands-on in the code.

  • Driving bi-weekly sprint planning and retros, contributing to the engineering roadmap, joining our daily 30-min remote stand-up at 7:30am PST (our only mandatory meeting), and taking part in the on-call rotation.

Example projects you could own:

  • Rearchitecting our patient data consolidation pipeline (deduplication, normalization, hydration) to handle 100x today's volume without 100x the cost.

  • Building pipelines that deliver clinical data directly into customers' data warehouses, reliably and at scale.

  • Designing the ingestion path for customers pushing large volumes of their own data into the platform.

  • Building document-processing pipelines that extract structured data from PDFs, images, and free text to feed ML models.

REQUIREMENTS

  • 8+ years of engineering experience, with significant depth building, operating, and scaling data platforms processing terabytes of data and millions-to-billions of events a day.

  • You've designed data architectures end-to-end — ingestion, storage, processing, warehousing, serving — and owned the tradeoffs (cost, latency, correctness, operability) at each layer.

  • Deep experience with modern, cloud-native data stacks: e.g., Spark, open table formats (Parquet, Iceberg, Delta) on S3, warehouses (Snowflake, BigQuery, Redshift), dbt, orchestration (Airflow, Dagster), and streaming (Kafka, Kinesis). Breadth matters — you'll be picking our stack.

  • Strong software engineering fundamentals — you write production code (we're a TypeScript shop, with Python in data/ML workflows), not just orchestration configs.

  • Experience mentoring or guiding other engineers — through code reviews, pairing, design feedback, or onboarding.

  • Located in San Francisco / Bay Area, or willing to relocate.

  • Bonus:

    • Experience leading engineers.

    • Experience building or supporting ML/data science workflows (feature pipelines, model inputs/outputs, unstructured data extraction).

    • Healthcare standards/technologies: FHIR, HIE, IHE, EHR/EMR, NPI, TEFCA, ADT, HL7, HEDIS, RAF, SNOMED, LOINC, ICD-10, etc.

BENEFITS

  • Competitive equity + compensation package 🚀

  • Full family Platinum health insurance, dental, and vision coverage 🦷

  • 401(k) retirement plan + matching 💰

  • Flexible work from home or in-office 🏢

  • Healthy lunches are complimentary when working in-office (and breakfast + dinners as needed) 🍏

  • Quarterly company off-sites with the team ⛷️

  • MacBook provided by us 💻

  • Unlimited PTO (we work hard, but trust you to take time you need to be at your best) 🧘‍♂️

OUR TECH

Core business logic in Node.js and TypeScript, with Python in data and ML workflows. AWS across the board (ECS, Lambda, SQS, SNS, Batch, etc.), infrastructure as code with CDK. Data lives in S3, PostgreSQL/Aurora, DynamoDB, Snowflake, and our FHIR server — with Athena for querying S3 and SageMaker for ML. Our data platform is still early: you'll shape what we adopt next, picking the best tool for the job rather than the trendiest one.

Metriport provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

You've read the whole posting — now see how you match it.