Senior Principal Software Engineer, Machine Learning
- Company
- Toast
- Location
- Remote
- Work type
- Full Time
- Posted
- 2026-08-10
Job description
A day in the life (Responsibilities):
Own technical direction of the ML Platform — feature store, model hosting and serving, experimentation, training infrastructure — driving architectural decisions around scalability, reliability, latency, and cost
Lead design and delivery of large-scope platform initiatives from conception through production, coordinating across ML, data, and infrastructure teams
Identify and resolve systemic technical challenges: online/offline feature parity, model deployment friction, experimentation velocity, GPU utilization, cross-team dependencies
Set and maintain a high engineering quality bar through hands-on code contributions, design reviews, and mentorship of platform and ML-adjacent engineers
Partner with ML engineering, data science, product, and platform leadership to translate ML strategy into technical roadmaps
Define the paved paths ML teams use to ship models safely — from feature registration through canary rollout, monitoring, and rollback
Leverage AI-augmented development tools to increase development velocity and code quality
What you'll need to thrive (Requirements):
10+ years delivering complex backend or infrastructure systems at scale
Direct experience building or operating core ML infrastructure — feature stores, model serving, experimentation platforms, training orchestration, or equivalent
Mastery of a modern backend language, ideally Java or Kotlin
Deep proficiency with distributed systems concepts: consistency, latency, throughput, fault tolerance, and observability
Strong understanding of data modeling, query languages, and the online/offline data patterns that underpin ML systems
Demonstrated technical leadership, with ability to drive cross-team alignment and influence engineering, product, and business stakeholders
Bachelor's degree in Computer Science or a related field, or equivalent practical experience
What will help you Stand Out (Nice to Haves):
Hands-on experience with open-source or commercial ML platform components (e.g. Tecton, MLflow, SageMaker, Databricks)
Experience building or operating experimentation / A-B testing platforms at scale
Familiarity with real-time streaming systems (Kafka, Flink, Spark Streaming) and their use in feature computation
Experience serving LLMs or deep-learning models in production, including GPU capacity planning and inference optimization
Prior work supporting internal-developer-facing platforms with a product mindset
Skills Required
10+ years delivering complex backend or infrastructure systems at scale
Direct experience building or operating core ML infrastructure (feature stores, model serving, experimentation platforms, training orchestration)
Mastery of a modern backend language, ideally Java or Kotlin
Deep proficiency with distributed systems concepts: consistency, latency, throughput, fault tolerance, and observability
Strong understanding of data modeling, query languages, and online/offline data patterns for ML systems
Demonstrated technical leadership with ability to drive cross-team alignment and influence stakeholders
Bachelor's degree in Computer Science or related field, or equivalent practical experience
Hands-on experience with open-source or commercial ML platform components (e.g., Tecton, MLflow, SageMaker, Databricks)
Experience building or operating experimentation / A-B testing platforms at scale
Familiarity with real-time streaming systems (Kafka, Flink, Spark Streaming) for feature computation
Experience serving LLMs or deep-learning models in production, including GPU capacity planning and inference optimization
Prior work supporting internal developer-facing platforms with a product mindset