← Back to jobs

Machine Learning Engineer

Location
New York City, NY
Work type
Full Time
Posted
2026-08-05

Job description

The Role
What You’ll Own
You’ll own the models inside Kepler’s AI research platform: which model runs each task, when a fine-tuned model beats a frontier one, and the training, evaluation, and extraction systems that make every workflow powerful. Model-agnostic by design doesn’t mean the model doesn’t matter. It means model choice is a permanent engineering problem, and it’s yours. The models you choose and tune sit inside a product financial professionals rely on for million-dollar decisions.

This role is for engineers who want to build foundational technology at the intersection of AI and finance, where your code directly impacts how clients make critical business decisions.

In the first few weeks you might:

Fine-tune a small model on a high-volume extraction task (footnote tables in 10-Ks, IR decks) and show it beats the frontier model we use today on accuracy, cost, and latency.
Build an eval harness that scores agent research runs end to end (does every number trace, does every citation resolve) and wire it into CI so regressions get caught before analysts see them.
Redesign model routing across a workflow: a frontier model where the reasoning is hard, cheaper or fine-tuned models for high-volume extraction and verification steps, with evals proving nothing got worse.
Take a workflow that succeeds 80% of the time and systematically find the other 20%: better tools, tighter verification rules, different context, a fine-tune, or a different model entirely.
In the longer term, you’ll be given ownership of whole functional areas, from extending our platform to a new industry to leading new architecture as our infrastructure scales.

You’ll consistently own systems end-to-end. In a small team, there’s nobody to hand things off to.

How We Work
We’re a close team, working together in an office in New York. We use AI tools heavily – Cursor, Claude Code, whatever makes us faster. Fluency is assumed. Our users are analysts at firms where a wrong number costs real money. The feedback loop on what you ship is hours, not quarters.

The pace is startup-fast but the engineering bar is high. We care about getting things right, not just getting things out. If you’ve worked somewhere that moves fast but ships broken software, this is different. If you’ve worked somewhere that’s rigorous but slow, this is also different.

The team has strong backgrounds and low ego. We expect everyone to roll up their sleeves and handle the unglamorous problems: the weird regressions, the subtle bugs, the last minute debugging session before a demo. We move as a team, not as a collection of individuals.

Who You Are
You’ve shipped production systems and you care about whether they’re correct – not just whether they work on the happy path. You think about failure modes before someone asks you to.

You’re comfortable in a codebase you didn’t write, moving between a fine-tuning run and the orchestrator code that serves the result in the same day. You’re drawn to early-stage not for the title but because you want your work visible in the product, not abstracted behind three layers of management.

From the technical side:

5+ years building production software. No upper limit, comp scales with experience.
You’ve shipped ML systems to production (fine-tuning, agents, retrieval, structured extraction) and you know what breaks between a demo and a product.
You treat evals as engineering: you build the measurement before the feature, and you don’t call something better until the numbers say so.
Strong general engineer, whatever your path into ML (research, ML infra, product). Our backend is Rust, but we don’t require Rust experience. We believe strong engineering fundamentals and experience in other languages is what matters.
You’re a quick learner and are as comfortable in a codebase you wrote as one you’re reading for the first time.
From the personal side:

You care what the analyst does with what you shipped, not whether the code was clever.
You’d rather fix something than file a ticket about it.
You’ll tell someone their design has a flaw before the PR goes in, not after.
You communicate before it’s a problem, and when a teammate needs something from you, they don’t have to ask twice.
You know what it feels like when the plan changes twice in a day and the work still has to ship.

Original source