Staff Data Engineer
- Location
- New York City, NY
- Work type
- Full Time
- Posted
- 2026-08-03
Job description
Core Responsibilities
Own ingestion pipelines end-to-end. Design and implement reliable, observable pipelines from operational sources into the Bronze layer using CDC patterns, Autoloader/Lakeflow, and batch ingestion.
Own the transformation layer. Build and maintain dbt models in a medallion architecture, set up testing and alerting, document data assets in Unity Catalog, and ensure sensitive data is tagged and access-controlled.
Drive GCP to AWS/Databricks migration. Take ownership of active pipeline cutovers from Airflow, Cloud Run, Cloud Functions, and BigQuery to Databricks on AWS. Validate parity, coordinate with stakeholders, and decommission legacy components without disrupting business-critical reporting.
Build and maintain reverse ETL and operational integrations. Sync curated data back into Salesforce, NetSuite, and other operational systems via Mulesoft.
Architect before building. Write ADRs for non-trivial design decisions. Define patterns the team follows for ingestion, transformation, and data quality. Make deliberate build-vs-managed tradeoffs and document the reasoning.
Translate business needs into data requirements. Partner with Sales, Finance, Marketing, Clinical, and Product teams to understand their data workflows and source systems. Convert that context into well-scoped pipeline and model work with clear acceptance criteria.
Qualifications
Baseline skills/experiences/attributes:
5+ years of data engineering experience, with clear evidence of seniority: owned complex domains, designed solutions from scratch, and made architectural decisions — not just executed tickets.
Strong hands-on dbt experience: models, tests, sources, macros, documentation, CI integration, and refactoring existing work.
Demonstrated experience with both streaming and batch ingestion patterns, including CDC pipelines, event-driven architectures, and scheduled bulk loads from operational sources.
Experience with Databricks (Delta Lake, Delta Live Tables) or a comparable lakehouse platform.
Solid Python for data engineering: PySpark, pipeline development, utilities, and custom tooling.
Hands-on experience integrating with operational data sources: CRM (Salesforce), ERP (NetSuite), payments (Stripe), or similar.
Ability to design before building: write the doc, define the interface, identify the failure modes, then implement.
Experience with data quality, testing, and pipeline observability: dbt tests, Great Expectations, alerting, SLA tracking.
Ideally, you also have these skills/experiences/attributes (but it’s ok if you don’t!):
Experience with Unity Catalog or an equivalent governance layer for column masking, row-level security, lineage, and sensitivity tags.
Familiarity with reverse ETL or integration platforms (Mulesoft, Boomi, or equivalent).
IaC experience at the data workload level (Terraform, Spacelift, or equivalents).
Experience in healthcare, life sciences, or another regulated environment (HIPAA, PII/PHI classification and handling).
Exposure to AI/ML platform patterns: vector search, RAG pipelines, or model serving data flows.