← Back to jobs

Agentic AI Engineer

Location
New York, NY
Work type
Full Time
Posted
2026-07-28

Job description

What You Will Do

You will design and build the specialist AI agents that form the core of the platform’s intelligence layer — each one reasoning over a different dimension of athlete performance, from strength and conditioning to recovery, readiness, and beyond.
You will develop the workflow engine that encodes domain scientist expertise into validated, versioned agent skills at scale, working directly with sport scientists to translate their judgment into calibration signals the system can act on reliably.
You will architect and build the decision intelligence layer that sits between agent outputs and practitioner delivery — combining confidence weighting, consequence classification, and human escalation logic to ensure every recommendation is defensible before it reaches a coach or performance director.
You will build toward the platform’s signature product experience: multiple specialist agents working together on a single complex question, synthesizing their findings into one coherent, traceable, calibrated recommendation in seconds.

WHAT YOU’LL NEED

5+ years in applied ML or AI engineering, with at least 2 years building production agentic AI systems — not chatbots, not RAG pipelines alone, but systems with memory, tool use, multi-step reasoning, and calibrated outputs
Deep experience with multi-agent frameworks and orchestration: dependency-aware routing, specialist agent composition, response synthesis across conflicting outputs
Hands-on experience with confidence calibration and evaluation frameworks for probabilistic systems — you understand Platt scaling, isotonic regression, and ECE, and you have built evaluation harnesses that run against full input distributions
Production RAG experience with reranking — you know that retrieval quality determines answer quality and have built pipelines that prove it
Experience fine-tuning or adapting foundation models for specific domains — knowledge injection, not general text
Strong Python, Golang; experience with LLM observability and drift detection in production

STRONGLY PREFERRED
Experience building knowledge acquisition workflows for domain-specific AI — annotation interfaces, version-controlled knowledge bases, review queues, regression testing against skill updates
Background working with domain scientists or clinical practitioners to encode expert knowledge into AI systems — you know how to translate judgment into calibration signals
Experience with human-in-the-loop architectures: escalation models, confidence thresholds, consequence classification
Familiarity with sports science, biomechanics, or performance data — understanding what “acute-to-chronic workload ratio” means matters in this role
Experience with causal or counterfactual reasoning in AI systems — not just pattern matching
Experience working with AWS (ECS, EC2, Lambda, SNS, SQS, etc), GraphQL, REST, gRPC, Postgres, Mongo

Original source