Node.Digital logo

Senior Data Engineer

Node.Digital

Washington, DCRemoteJobNo compensation foundTracked 3w agoSeen in employer's feed 2 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Washington, DCRemote
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Public Trust clearanceBachelor's degree

Job overview

Node.Digital is hiring a Senior Data Engineer. The Senior Data Engineer will provide authoritative expertise on data engineering methods, design and maintain data architecture and pipelines in Azure, migrate source data, normalize entities, establish quality controls, and produce documentation while staying current with emerging AI tooling.

Key focus areas include Provide authoritative expertise on data engineering methods and best practices, Design, implement, and maintain data architecture supporting products and end users, and Design, implement, and maintain ELT and ETL pipelines in Azure Synapse and Azure Machine Learning.

Important skills include Data Engineering Methods, Code First Development Approaches, Modern Pipeline Design Patterns, Data Architecture Design, Source Control, and ELT. Preferred (not required): Python, PySpark, Polars, and Reusable Modular Code Development.

Skills & qualifications

RequiredNice to have

Skills

Data Engineering MethodsCode First Development ApproachesModern Pipeline Design PatternsData Architecture DesignSource ControlELTETLAzure SynapseAzure Machine LearningSDK V1SDK V2Azure Data Lake StorageEntity Attribute NormalizationArchitecture ReviewPipeline ImprovementQuality ControlsError HandlingLogging MechanismsValidation ChecksIngestion OptimizationProcessing OptimizationStorage OptimizationParquetSelf Service Capabilities DevelopmentStandard Operating Procedures AuthoringData Architecture DocumentationData DictionariesEntity Relationship DiagramsPipeline Process MapsEnvironment MaintenanceAI ToolingAutomation EvaluationLanguage Model Assisted CapabilitiesSQLPythonPandasPySparkPolarsReusable Modular Code DevelopmentCLIREST APIsTerraformBicepCI/CDContinuous Delivery WorkflowsAI Coding AssistantsLarge Language Model Integration PatternsEntity ResolutionSelf Service Analytic Access

Qualifications

Public Trust ClearanceBachelor's Degree in Data Engineering, Computer Science, Data Science, Machine Learning, Mathematics or Related Field5 Years Applied Work Experience in Data Engineering, Computer Science, Data Science, Machine Learning, Mathematics or Related Field5+ Years Maintaining SQL Databases5+ Years Designing, Implementing, and Maintaining ELT and ETL Processes in Cloud Based Data Analytics Environments3+ Years Working in Azure Synapse and Azure Machine Learning With the Modern Data Stack3+ Years Manipulating Data in PythonDP-203Microsoft Certified Azure Data Engineer Associate

Benefits

Medical Insurance
Dental Insurance
401(k) Match
Paid Time Off

Full job description

Senior Data Engineer

Location: Herndon, VA (Remote Work)

Must have an Public Trust Clearance

KEY RESPONSIBILITIES

  • Provide authoritative expertise on data engineering methods and best practices, including code first development approaches and modern pipeline design patterns.

  • Design, implement, and maintain the data architecture that supports products and end users, with all assets managed under source control.

  • Design, implement, and maintain ELT and ETL pipelines for efficient processing of source data in Azure Synapse and Azure Machine Learning, using both SDK V1 and SDK V2.

  • Migrate source data identified by SBA OIG into Azure Data Lake Storage.

  • Normalize entity attributes such as addresses, phone numbers, and other common fields.

  • Review, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift.

  • Establish quality controls across all pipelines and introduce error handling, logging mechanisms, and validation checks.

  • Incorporate source control across all pipelines and analytics codebases so code can evolve iteratively without destabilizing the architecture.

  • Optimize ingestion, processing, and storage across a wide variety of datasets and data types, including modern columnar formats such as Parquet.

  • Develop self service capabilities that let SBA OIG analysts query and export data for investigations and audits.

  • Author robust standard operating procedures governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment.

  • Produce detailed documentation of the data architecture, including data dictionaries, entity relationship diagrams, and pipeline process maps.

  • Maintain and expand the environment with additional datasets and services on request, following a defined intake and testing process before production deployment.

  • Stay current with emerging AI tooling relevant to data engineering and contribute to exploratory work evaluating automation and language model assisted capabilities.

Requirements

Education

Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field. Alternatively, five years of applied work experience in any of the same fields.

  • 5 years - Maintaining SQL databases and conducting advanced operations in SQL and T-SQL.

  • 5 years - Designing, implementing, and maintaining ELT and ETL processes in cloud based data analytics environments.

  • 3 years - Working in Azure Synapse and Azure Machine Learning with the modern data stack. Certifications preferred, DP-203 or equivalent.

  • 3 years -Manipulating data in Python. Pandas is required. PySpark and Polars preferred. Experience developing reusable, modular code preferred.

PREFERRED QUALIFICATIONS

  • DP-203, Microsoft Certified Azure Data Engineer Associate, or an equivalent current certification.

  • Implementing pipelines and infrastructure using code first approaches: Python SDK, CLI, REST APIs, or infrastructure as code tooling such as Terraform or Bicep.

  • Implementing source control and continuous integration and delivery workflows for data assets.

  • Demonstrated familiarity with AI coding assistants and large language model integration patterns.

  • PySpark or Polars at production scale.

  • Entity resolution and attribute normalization across records with inconsistent addresses, names, and identifiers.

  • Building self service analytic access for non engineering users.

Benefits

We are proud to offer competitive compensation and benefits packages to include

  • Medical

  • Dental

  • Vision

  • Basic Life

  • Health Saving Account

  • 401K matching

  • Three weeks of PTO/Sick

  • 11 Paid Holidays

  • Pre-Approved Online Training

You've read the whole posting — now see how you match it.