Data Engineer
- Location
- Manhattan, NY
- Work type
- Full Time · Remote
- Posted
- 2026-08-05
Job description
The Data Engineer Role.
We’re looking for a Data Engineer who can build and scale the infrastructure powering our data platform, with a strong foundation in distributed systems and cloud-native tooling. You’ll design and operate the pipelines that move, process, and prepare multi-terabyte and streaming datasets — including audio and video — for Reality Defender’s detection models, and you’ll work closely with ML engineers and researchers to keep that data flowing reliably at scale. Responsibilities include:
Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows.
Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments.
Build and maintain distributed data processing jobs using frameworks such as Spark and Ray.
Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines.
Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies.
Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.
Who you are.
Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments.
Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets.
Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines.
Strong programming skills in Python and SQL; experience with Golang is a plus.
Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets.
Experience working with audio or video data at scale is a strong plus (e.g., transcoding, feature extraction, or preprocessing pipelines).
Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks).
Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines.
Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling).