Staff Software Engineer – Data
- Location
- New York City, NY
- Work type
- Full Time
- Posted
- 2026-08-05
Job description
What you’ll do
Building integrations that ingest and sync customer systems (EHRs, schedulers, warehouses, APIs)
Designing transformations that turn messy source data into one normalized model
Building and optimizing data pipelines that keep one clean profile per patient
Powering natural language query interfaces over healthcare data
Owning data modeling, query performance, and data freshness at scale
What we’re looking for
Built and operated production data platforms that ingest, process, and serve millions of events with high reliability
Designed scalable streaming and batch pipelines, data models, and ETL/ELT workflows for production systems
Possess deep expertise in SQL, distributed query optimization, and large-scale data processing
Have hands-on experience with modern data platforms such as Databricks, Snowflake, Delta Lake, Apache Iceberg, Spark, Kafka, or similar technologies
Designed event-driven architectures, change data capture (CDC), online serving systems, or reverse ETL pipelines
Built connector frameworks or ingestion platforms that integrate enterprise applications and third-party data sources
Balance performance, scalability, cost, and operational simplicity when designing distributed systems
Own production systems end-to-end, including architecture, implementation, monitoring, reliability, and incident response
Value simple, maintainable solutions, communicate directly, and maintain a high engineering bar with a low-ego, collaborative approach
You can work on site in New York City or San Francisco
Nice to have
Experience building data platforms in regulated or high-reliability industries such as healthcare, financial or services
Familiarity with healthcare data standards such as FHIR, HL7, or other clinical interoperability frameworks
Experience building data platforms that power production AI, machine learning, or agentic applications
Familiarity with modern lakehouse technologies such as Delta Lake, Apache Iceberg, or Unity Catalog
Experience with streaming platforms, change data capture (CDC), or event-driven architectures
Experience operating large-scale data platforms with a focus on reliability, observability, and cost efficiency