Head of AI Engineering at AIOS
- Company
- AIOS
- Location
- Remote US
- Work type
- Full Time
- Posted
- 2026-08-12
Job description
Being a Head of AI Engineering at AIOS
We are building a world-leading Applied AI team.
As Head of AI Engineering at AIOS, your fundamental role is to build the AIOS Agent SDK and make it the foundation for world-class agents across the company.
You are not joining to discover our first AI use case or build another chatbot.
Our customer-facing agent gathers context from across our product and customer history, retrieves the right knowledge, reasons through multi-step cases, and decides when to act, respond, escalate, or stand down. Our agentic workflows autonomously generate >$100k of revenue per day.
This existing harness will become the nucleus of the AIOS Agent SDK. You’ll separate its reusable foundations from its customer-support logic and turn them into a strongly opinionated internal platform.
We’re also building an AI clinical decision-support system. This helps clinicians evaluate patient eligibility, contraindications, dosing, and risk. It will be the second major system built on the SDK and, over time, a foundation for increasingly autonomous clinical decisions.
Once the SDK has proven itself through these two tools, it will become the default foundation for new agents across AIOS.
You’ll be the DRI for agent architecture, model strategy, evals, AI reliability, technical safety, provider relationships, and the shared runtime. You’ll make these decisions autonomously. You’ll ensure we use the best model for each job based on measured quality, reliability, speed, and cost.
This is a technical leadership role. You’ll lead by example as you grow the team.
You’ll ensure:
The AIOS Agent SDK exists and is running the show in production, and it is in exceptionally safe technical hands
Product engineers can build excellent agents without recreating context, tool, safety, eval, and observability infrastructure.
Our agents become more capable without becoming less predictable.
Major changes are supported by trustworthy evidence across quality, reliability, safety, latency, and cost.
Production failures continuously strengthen our evals, architecture, and models.
Our engineers actively seek your judgment and trust the direction you set.
AIOS is clearly an industry leader in applied AI for production healthcare systems.
Key responsibilities
Agent SDK: You’ll turn Jesse’s (customer support tool) existing harness into the strongly opinionated internal platform powering Jesse, Aegis (clinical support tool), and future AIOS agents. You’ll own its architecture, reusable primitives, supported extension points, developer experience, and integration with our existing infrastructure.
Jesse & Aegis: You’ll become the senior technical owner of Jesse and work closely with the engineers and Clinical Product team building Aegis. You’ll improve both systems while extracting the shared foundations they need across context, retrieval, memory, orchestration, tools, state, and escalation.
Evals & Experimentation: You’ll build trustworthy benchmarks using deterministic checks, simulations, model-based graders, human judgment, and production outcomes. You’ll establish the path from offline evaluation to controlled production experiments so major changes ship with evidence.
Production Learning Loop: You’ll turn traces, poor resolutions, escalations, incidents, tool failures, and successful outcomes into better evals, stronger architecture, improved models, and permanent platform capabilities.
Safety & Compliance: You’ll make consequential agent actions safe through authorization, validation, idempotency, auditability, recovery, and human handoff. You’ll encode compliance, privacy, security, and regional requirements into the platform wherever possible.
Models & Economics: You’ll own model selection, routing, fallbacks, caching, and our ~$200k monthly model spend. When the evidence supports it, you’ll lead the data preparation, fine-tuning, evaluation, and AIOS-controlled deployment of specialized open-weight models.
Reliability: You’ll own the shared runtime in production, including tracing, observability, testing, provider resilience, capacity, and incident response. You’ll be the senior engineering DRI when an AI system behaves unsafely, quality regresses, or the platform fails.
Technical Leadership: You’ll set AIOS’s AI architecture and strategy in close partnership with the VP of Engineering. You’ll make the final call on major technical decisions, guide engineers across product pods, and remain hands-on by writing production code and personally building the most important foundations.
Build the Team: You’ll inherit one engineer and build the Applied AI team to approximately five exceptional people during your first year. You’ll own our technical relationships with leading model providers and represent AIOS externally where doing so strengthens our work.
Need to have
Experience: You have 8+ years of software engineering experience and remain an active production contributor.
Education: You have at least a bachelor’s degree in Computer Science, Machine Learning, or a closely related technical field.
Production Agents: You have personally built and shipped an exceptional agentic system used by real customers. It did more than answer questions: it reasoned across multiple steps, used tools, changed state, and operated under real production constraints.
Agent Architecture: You can reason deeply about harnesses, orchestration, context construction, retrieval, memory, state, tool design, structured workflows, and error recovery.
Evals: You’ve built or meaningfully owned evaluation systems for probabilistic products. You understand dataset construction, evaluator design, simulations, regression detection, noisy metrics, and the relationship between offline performance and production outcomes.
Software Engineering: You have strong systems-engineering fundamentals. You can reason about APIs, distributed systems, concurrency, queues, databases, observability, failure modes, and production reliability.
Consequential Actions: You know how to let an agent act safely. You have strong judgment around authorization, validation, idempotency, state transitions, auditability, recovery, and escalation.
Model Judgement: You understand the capabilities and limitations of current frontier and open-weight models. You know when the model is the problem and when the real problem is context, tools, data, orchestration, or evaluation.
Open-Weight Models: You have enough technical depth to lead the fine-tuning and AIOS-controlled deployment of open-weight models when the evidence supports doing so. Prior production deployment is not required.
Leadership: You have successfully led and managed a small technical engineering team. You set a clear direction, raise the quality bar, develop strong engineers, and address underperformance.
Technical Authority: Strong engineers trust your judgment. You can make difficult decisions, explain the trade-offs clearly, and push back without hesitation when a proposed approach is unsound.
Communication: You can explain difficult technical ideas to engineers, product leaders, clinicians, and executives without flattening the important details.
Independence: You create clarity in ambiguous environments and make high-quality decisions without hand-holding.
Builder: You still write production code. You lead from inside the work rather than managing it from a distance.
Ownership: When quality drops, costs spike, tools fail, or providers degrade, you take responsibility for reaching the outcome rather than identifying whose component was technically at fault.
Nice to have
Agent Platforms: You’ve built runtimes, SDKs, harnesses, tool layers, evaluation platforms, or shared AI infrastructure used by other engineers.
Customer Agents: You’ve built high-volume customer-service, commerce, or transactional agents operating across complex, multi-step customer journeys.
High-Stakes Systems: You’ve worked on healthcare, financial, insurance, or other systems where correctness, traceability, and careful rollout matter.
Model Adaptation: You’ve fine-tuned, distilled, evaluated, or deployed an open-weight model for a specific production workflow.
Long-Term Memory: You’ve built durable memory, context compression, personalization, or agents operating across sessions and extended periods.
Real-Time Systems: You’ve worked on voice agents, streaming systems, or other latency-sensitive AI experiences.
Provider Relationships: You’ve worked directly with frontier model providers on evaluations, technical issues, capacity, pricing, or early access.
Talent: You have a strong nose for exceptional AI engineers and know how to create an environment in which they do their best work.
Research Fluency: You can translate relevant research into reliable production systems without confusing novelty with progress.
Figure It Out: You can move from debugging a production trace, to redesigning an eval, to reviewing an agent abstraction, to handling a provider incident.