← Back to jobs

Staff ML Platform Engineer

Company
Datavant
Location
Remote - United States
Work type
Full Time
Posted
2026-09-29

Job description

What We’re Looking For

We are looking for a Staff ML Platform Engineer to join our Data & ML Platform organization on the ML Platform team. We own the paved road that takes a model from notebook to production without rebuilding it each time: training and inference across SageMaker and Databricks, self-service LLM-endpoint serving, and the pipelines behind Datavant’s clinical-AI products. Data Science owns the models and their quality; we own the pipelines, serving infrastructure, and operational guardrails that let those models run safely against regulated health data.

As a Staff ML Platform Engineer, you will be a technical leader across the ML Platform: setting direction for the paved road, owning the hardest architectural problems, and moving the team from bespoke plumbing toward a coherent platform that Data Science teams can self-serve. AI fluency is a baseline expectation. You should already be using Claude Code, Cursor, Copilot, or equivalent tools as a core part of your daily engineering workflow, have opinions about how they make a team faster, and know how to apply them responsibly when PHI and other sensitive data are in scope.

What You Will Do

Set technical direction across ML training, serving, and observability, and be the final escalation point for the most elusive infrastructure problems (GPU capacity, Spark tuning, production incidents)
Own and evolve our paved-road framework (the shared CI/CD spine, model-workflow scaffolding, and Databricks Asset Bundles) so Data Science teams can go from a config file to a production workflow without bespoke plumbing
Lead architecture for LLM-endpoint serving across managed providers (Databricks, AWS, Snowflake) and self-hosted deployments, covering latency, cost, caching, evaluation, and PHI-safe routing
Own the standards and tooling for MLflow, model registry, training image supply chain, and observability across training and inference
Partner closely with your Data & ML Platform teammates to present a cohesive ML platform to the Data Science, App Dev, and Operations teams at Datavant
Serve as a key technical input to vendor and platform selection decisions across model providers, ML tooling, and observability
Mentor senior engineers on the team, provide technical guidance to platform consumers, and stay hands-on writing high-leverage code and Infrastructure-as-Code alongside your teammates
What You Need to Succeed

10+ years of software engineering experience, with 3+ years designing, evolving, and operating enterprise-scale ML platforms in production
Strong technical judgment under ambiguity and a track record of setting standards, influencing peers, and raising the bar across teams
Hands-on production experience with Databricks and/or Amazon SageMaker, MLflow (or an equivalent tracking + registry system), and at least one core ML framework (PyTorch, TensorFlow, or similar)
Fluency in Java (or a JVM equivalent) and Python, with real depth in Apache Spark for large-scale data and distributed compute
Real depth in AWS: networking, IAM, GPU compute, and the storage and messaging services this role touches, with the judgment to know what to reach for and when
Fluency with Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads
Direct experience serving LLMs in production, including cost management, evaluation harnesses, and safe handling of sensitive prompts and outputs
AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster
Clear written and verbal communication, especially in async, remote settings

What Helps You Stand Out

Prior technical leadership on a healthcare or regulated-industry ML platform (HIPAA, HITRUST, SOC 2, or equivalent)
Direct experience with Databricks Asset Bundles, Unity Catalog, and open table formats (Iceberg, Delta)
Production experience with specialized inference pipelines beyond generic model serving (e.g., document understanding, computer vision, or streaming inference)
GPU capacity planning at strategic, tactical, and operational horizons
Experience with real-time and event-driven inference and streaming platforms (Kafka, Kinesis)
Evaluation and red-teaming experience for clinical or safety-sensitive AI
Background contributing to open-source ML infrastructure or publishing on production ML systems

Original source