← Back to jobs

Research Engineer, LangSmith Engine

Company
LangChain
Location
New York, NY
Work type
Full Time
Posted
2026-08-16

Job description

About the Role:
We’re looking for an experienced research engineer to help make the Engine agent more capable and more efficient.

You’ll study real agent failures, build benchmarks that capture what matters, run experiments to improve performance, and turn successful ideas into production. This may include prompting and agent-harness improvements, model selection, fine-tuning and post-training custom models. The focus is on measurable improvements to the overall agent.

This role also requires a understanding of production engineering and system-level tradeoffs. Engine is a production system, so improving an agent is not just about maximizing benchmark performance—it also means understanding the impact on cost, latency, reliability, and scalability. You’ll work in the same team with production engineers to design, test, and ship improvements that work reliably in real-world environments.

Location: SF and NYC

What You’ll Do:
Build and maintain benchmarks and evaluations that measure the quality and efficiency of Engine agents on real-world tasks.

Design and run experiments to improve agent performance across models, prompting, context, tools, orchestration, and agent strategies.

Explore and implement post-training and fine-tuning techniques when they can meaningfully improve agent capabilities, quality, or cost.

Turn successful experiments into production improvements, working closely with engineers and researchers to measure impact and prevent regressions.

Help define the ML roadmap and technical direction for improving Engine agents, and mentor other engineers through strong technical leadership.

What You’ll Bring:
4+ years of experience in ML/AI research, or a closely related field.

Master’s or Ph.D. in a relevant scientific field.

Hands-on experience working with LLMs and AI agents, including analyzing model behavior and improving real-world performance

Strong experience designing benchmarks, evaluations, and experiments for AI/ML systems; you know how to tell whether a change actually made an agent better.

Strong software engineering skills, with a track record of taking ideas from research prototype to measurable production impact.

You have maximum agency and strong research judgment: you can identify high-impact problems, work through ambiguity, move quickly, and communicate your findings clearly.

Nice to Have:
Ph.D. in Machine Learning, Computer Science or Physics.

Hands on experience with LLM-as-a-judge, automated graders, synthetic data generation, or human evaluation.

Hands on experience with reinforcement learning, preference optimization, SFT, RLHF/RLAIF, or other post-training techniques for LLMs.

Experience optimizing LLM agents for cost, latency, or task efficiency on productions

Experience with model serving, inference optimization, distributed systems, or GPU infrastructure.

Original source