Staff Machine Learning Engineer, ML Platform
- Company
- Braze
- Location
- New York City, NY
- Work type
- Full Time
- Posted
- 2026-09-17
Job description
As the Staff Engineer on the team, you will:
Identify and drive the transformative initiatives that change how the team runs ML in production, whether that's replatforming our queueing and orchestration, overhauling deployment and cloud identity, or retiring a generation of infrastructure
Build and ship at high velocity. Staff at Braze is a hands-on delivery role; you carry the most complex infrastructure initiatives yourself from design through production. Current examples include multi-region model serving fleets, the pipelines that keep hundreds of customer-specific models healthy, and the CI and deployment tooling that moves it all safely
Own the platform's technical vision and production quality bar. Set direction for how models are trained, deployed, served, and observed; lead incident response for ML systems; and drive the reliability and cost work that keeps the platform efficient at scale
Drive initiatives that span teams. Our platform builds on shared infrastructure, deployment tooling, and data systems owned with partner teams, and you carry the technical relationships with those teams
Raise the team's engineering quality through design review, code review, and production readiness for ML systems, and mentor other senior engineers and data scientists
Connect technical decisions to customer and business outcomes, and represent the team's technical perspective to product and engineering leadership
WHO YOU ARE
8+ years building and operating distributed systems in production, with depth in deployment and operations. You have designed services for scale and reliability, owned CI/CD and infrastructure as code, and run what you built under production load
Hands-on experience with ML workloads in production. Training pipelines, model serving, feature systems, or ML platform tooling all count; deep modeling experience is a plus rather than a requirement
A technical leader who has owned direction for a team, led multi-quarter initiatives across team boundaries, and grown senior engineers, all while keeping a high personal output
Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management, networking, and the cost profile of what you run
An effective communicator, both verbal and written, whose designs and recommendations build consensus and drive forward decision making
Bonus:
Queueing and orchestration systems such as Celery, RabbitMQ, Kafka, or Ray
ML platform tooling such as MLflow or another model registry, feature stores, or ML observability
Experience in our stack (Python, Ruby on Rails, MongoDB, Redis, Kubernetes)
Operating under compliance regimes such as SOX or HIPAA
Customer engagement, personalization, or marketing technology domain experience