Engineering Manager – Machine Learning
- Location
- Toronto, CA
- Work type
- Full Time · Remote
- Posted
- 2026-07-20
Job description
In This Role You Will:
Enable AI/ML, LLM, and Agentic Systems teams for scale – The ML infrastructure team is responsible for building and operating platforms that allow data scientists and ML engineers to train, deploy, and monitor models across Recursion’s massive datasets. With billions of compounds, 30+ petabytes of experimental data, and complex deep learning workloads, your team enables everything from automated compound screening models to clinical trial prediction systems. You will work closely with researchers and ML engineers to understand their infrastructure needs and build scalable solutions for model development, training, and deployment.
Act as a mentor, coach, and sponsor – You will share your technical, leadership and managerial skills in MLOps, distributed computing, and infrastructure engineering, delivering impact, learning, and growth across teams at Recursion. We believe that the best work comes from working across organizational boundaries and you will have opportunities to partner with ML research, platform engineering, and business teams.
Enable a model-driven culture – Machine learning is at the core of everything we do. You will work with stakeholders across the business to ensure our ML infrastructure supports rapid experimentation, reliable model deployment, and continuous improvement. Problems you will work on could range from optimizing GPU cluster utilization to implementing Agentic orchestration and establishing company-wide MLOps standards
The Experience You Will Need:
Experience in a hands-on technical role as a tech lead or a manager with a focus on infrastructure, MLOps and distributed systems. Excitement for deeply engaging in technical details with your team around machine learning, orchestration and agentic systems.
A people-first mindset. We deliver in a way that prioritizes supporting our coworkers in their growth and experience and understand how Conway’s Law shapes our ML system outcomes.
Demonstrated past record of learning from and teaching peers in areas of ML infrastructure, model deployment, distributed compute, GPU optimization, and MLOps system architecture
Excitement to learn parts of our ML tech stack that you might not already know. Our current ML infrastructure includes: Python, PyTorch, Docker, Kubernetes, Ray, Weights & Biases, Prefect, BigQuery, Postgres, GCP, CUDA, and various model serving frameworks.
Fluency in life sciences or drug discovery is a plus but not required to be considered.