Polygon.io logo

Data Acquisition Engineer

Polygon.io

United StatesRemoteFull-timeNo compensation foundPosted 7mo agoVerified open 5 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
United StatesRemote
Schedule
Full-time
Work Authorization
Not specified

Job overview

Polygon.io is hiring a Data Acquisition Engineer. Polygon.io is seeking a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets for their financial data platform. The role involves transforming various inputs into clean, well-organized, analysis-ready assets, handling the full technical onboarding lifecycle from raw data ingestion to producing high-quality structured outputs. This position is ideal for someone who enjoys bringing order to messy datasets and building reliable ingestion pipelines.

Key focus areas include Own the technical onboarding of external datasets, including parsing raw files, transforming fields, and producing clean structured outputs., Write parsing and transformation logic in Python and SQL to handle diverse file formats (CSV, JSON, XML, HTML, XBRL, PDF, etc.)., and Develop reproducible ETL/ELT workflows that clean, normalize, validate, and structure incoming datasets..

Successful candidates bring Direct Experience With Financial Datasets, Experience Designing Etl ElT Workflows, and Experience Extracting Structured Information From Pdfs. Important skills include Python, SQL, S3-Compatible Object Storage, Parquet, Columnar Storage Formats, and ETL/ELT Workflows. Preferred (not required): DataFusion, DuckDB, Modern Analytical Engines, and Regulatory Filings Datasets.

Skills & qualifications

RequiredNice to have

Skills

PythonSQLS3-Compatible Object StorageParquetColumnar Storage FormatsETL/ELT WorkflowsParsing Logic DesignClear Written CommunicationDataFusionDuckDBModern Analytical EnginesRegulatory Filings DatasetsFinancial Reference Data DatasetsAlternative Data DatasetsExtracting Structured Information From PDFsSchema ValidationMetadata ManagementData-Quality Frameworks

Qualifications

Direct Experience With Financial DatasetsProven Track Record Onboarding or Structuring Large, Complex Financial or Finance-Adjacent Datasets

Full job description

Data Acquisition Engineer

Product

USA

Engineering

We are looking for a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets used across our financial data platform. You will work with a wide range of inputs and turn them into clean, well-organized, analysis-ready assets.

Apply Now About the Role We are looking for a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets used across our financial data platform. You will work with a wide range of inputs—including public datasets, financial reference data, and alternative data—and turn them into clean, well-organized, analysis-ready assets.

You will take datasets that have been identified for integration and handle the full technical onboarding lifecycle: ingesting raw data, designing parsing logic, developing cleaning and normalization workflows, implementing validation checks, and producing high-quality structured outputs. If you enjoy bringing order to messy, heterogeneous datasets and building reliable ingestion pipelines that support financial research and analytics, this role is for you.

Responsibilities

  • Own the technical onboarding of external datasets, including parsing raw files, transforming fields, and producing clean structured outputs.
  • Write parsing and transformation logic in Python and SQL to handle diverse file formats (CSV, JSON, XML, HTML, XBRL, PDF, etc.).
  • Develop reproducible ETL/ELT workflows that clean, normalize, validate, and structure incoming datasets.
  • Manage data storage and processing workflows using S3-compatible object storage systems.
  • Produce efficient, analytics-ready Parquet datasets, using appropriate partitioning and metadata conventions.
  • Implement data-quality checks to detect anomalies, schema drift, missing fields, or unexpected changes in incoming data.
  • Troubleshoot and resolve inconsistencies through systematic, transparent cleaning and transformation rules.
  • Collaborate with internal data and research teams to understand dataset characteristics, quirks, semantics, and intended uses.
  • Provide light technical input during dataset evaluation, offering insight into ingest feasibility and transformation complexity.
  • Write clear documentation describing dataset structure, parsing assumptions, transformation logic, and known limitations. Skills & Qualifications
  • Direct experience working with financial datasets at a hedge fund, financial institution, or major financial data provider.
  • Proven track record onboarding or structuring large, complex financial or finance-adjacent datasets, such as market data, fundamentals, regulatory data, reference data, or alternative data.
  • Strong proficiency in Python for parsing, cleaning, and transformation workflows.
  • Strong SQL skills for exploration, validation, and modeling.
  • Hands-on experience working with S3-compatible object storage for large dataset management.
  • Proficiency with Parquet and other columnar storage formats, including partitioning strategies for performance and scale.
  • Experience designing ETL/ELT workflows that are reproducible, maintainable, and resilient to upstream dataset changes.
  • Ability to interpret messy or loosely documented datasets and design stable parsing logic.
  • Clear written communication skills for documenting processes, assumptions, and dataset behavior.
  • Experience with DataFusion, DuckDB, or other modern analytical engines is a plus.
  • Exposure to datasets such as regulatory filings, financial reference data, or alternative data is a plus.
  • Experience extracting structured information from PDFs or other irregular data sources is a plus.
  • Familiarity with schema validation, metadata management, or data-quality frameworks is a plus. About Massive At Massive we are on a mission to help developers build the future of fintech. We are committed to democratizing access to the world's financial market data and enabling developers to build the future of fintech. Join us and be part of a team that is revolutionizing the way we interact with money and value.

If you are a passionate problem-solver who thrives in a dynamic and innovative environment, we encourage you to apply!

You've read the whole posting — now see how you match it.