Senior Backend Engineer (ML)
- Location
- New York City, NY
- Work type
- Full Time
- Posted
- 2026-08-03
Job description
Role Description
We’re looking for a founding Sr. ML Infrastructure Engineer with a strong background in distributed systems and building pipelines that scale. In this role, you’ll own the infrastructure that powers Tennr’s AI-driven healthcare platform – the training, inference, and data pipelines that let our models handle growing traffic and an expanding product surface.
Our ML team builds in-house, proprietary VLMs, LLMs, and other models purpose-built for hard problems in healthcare. You don’t need deep ML experience to thrive here – what matters is a strong system design foundation and the interest to grow into the ML-side. If you think in systems, care about reliability, and want to expand into ML infrastructure, this is a rare chance to build foundational systems from the ground up.
Responsibilities
Architect, build, and scale the cloud infrastructure behind our ML training, inference, and data pipelines.
Design resilient systems for model deployment, evaluation, and monitoring that stay reliable as traffic grows.
Own observability across the stack – logging, metrics, tracing, and alerting.
Troubleshoot production issues and continuously improve performance and efficiency.
Collaborate with ML engineers, backend engineers, and cross-functional teams to integrate models cleanly with data pipelines and products.
Candidate Qualifications
4+ years building and scaling infrastructure in production-distributed systems, cloud platforms, or data engineering.
Strong backend software engineering fundamentals, with proficiency in Python and TypeScript.
Hands-on experience with AWS, and PostgreSQL
Solid grasp of observability, reliability, and production incident response.
Comfortable with ambiguity and high ownership; you move fast and drive projects from idea to production in a startup environment.
Interested in growing into ML infrastructure – prior ML ops/infra experience is not required.
Nice to have: exposure to inference engines (vLLM, SGLang, TensorRT), or k8s.