← Back to jobs

Senior Data Engineer

Company
Chicory
Location
Remote - United States
Work type
Full Time · Hybrid
Posted
2026-10-08

Job description

Responsibilities

Map how ad requests, bids, impressions, clicks, campaign metadata, spend, revenue, publisher payments, partner files, and product events move from their source into BigQuery and Superset
Choose and implement the new BigQuery structure, including raw, staging, core, and reporting datasets; table grain and keys; naming; partitioning and clustering; incremental loads; and ownership
Build and maintain batch and streaming ETL and ELT pipelines with Python, SQL, Apache Beam, and Dataflow using data from Pub/Sub, Cloud Storage, APIs, databases, and files
Build and maintain Cloud Composer and Apache Airflow DAGs with clear dependencies, useful retries and alerts, and tasks that can be rerun safely
Choose and implement a version-controlled SQL transformation workflow using Dataform, dbt, or an equivalent tool, then move important calculations out of scheduled queries and Superset dashboards
Build the BigQuery tables and data marts used for campaign delivery, spend and revenue reconciliation, publisher payments, partner reporting, product analytics, and leadership reporting
Investigate mismatches between source systems, partner reports, BigQuery, and Superset; correct the logic, repair affected data, and document why the numbers changed
Add tests and alerts for late or missing data, duplicate records, unexpected schema changes, broken joins, and totals that do not reconcile
Maintain the Superset datasets that power business reporting, review expensive or confusing queries, and make sure shared metrics come from the same maintained model
Review BigQuery usage and cost, then improve the tables, queries, schedules, storage, or retention rules responsible for avoidable spend
Manage data access with GCP IAM, limit access to sensitive fields, and document retention or deletion rules where they apply
Use Git, automated tests, code review, deployment pipelines, and infrastructure as code so changes can be reviewed, repeated, and rolled back
Work directly with Engineering, Campaign Management, Finance, Product, Sales, and leadership to define calculations, verify new tables, and decide when old reports can be retired
When a pipeline fails or a report looks wrong, lead the investigation and recovery. Leave runbooks, diagrams, model descriptions, and metric definitions that another engineer can follow
Build upon team and company culture and act as a champion of the Chicory Principles

Qualifications

Demonstrated success designing, rearchitecting, and operating a production analytical data platform that includes a data lake or durable raw layer, a cloud data warehouse, governed data models, and business-facing data marts
Deep hands-on knowledge of BigQuery, including schema and table design, partitioning and clustering, query plans and performance, workload patterns, permissions, retention, and cost management
Strong Python and advanced SQL skills, including the ability to write maintainable, tested production code and review work across ingestion, transformation, orchestration, and data-serving layers
Hands-on production experience building batch and streaming systems with Apache Beam and Dataflow and with Apache Airflow and Cloud Composer, including retries, idempotency, dependency management, duplicate and late-arriving events, schema evolution, replay, backfills, monitoring, and failure recovery
Strong data-modeling judgment across dimensional, event, and domain models, with experience creating stable facts, dimensions, slowly changing dimensions, semantic definitions, and reusable analytical datasets
Track record of establishing automated data quality, reconciliation, observability, lineage, alerting, service-level objectives, incident response, and operational runbooks for critical data products
Experience enabling BI and self-service analytics through Superset or a comparable platform, including curated datasets, governed metrics, access patterns, dashboard performance, and stakeholder education
Applied understanding of data governance, least-privilege access, privacy and sensitive-data handling, auditability, retention, and the controls required for financial or revenue-impacting data
Experience with software-engineering practices for data systems, including Git, automated testing, CI/CD, code review, infrastructure as code such as Terraform, and clear technical documentation
Demonstrated ability to operate as an autonomous senior technical owner: assess an unfamiliar environment, make architectural decisions, sequence migrations, resolve ambiguity, lead incidents, and communicate tradeoffs with both engineers and non-technical stakeholders

Original source