← Back to jobs

Software Engineer, Verification Fleet

Location
Los Angeles, CA
Work type
Full Time · Remote · Remote
Posted
2026-08-06

Job description

What You Will Own
The robot fleet, end to end. The browser automation that drives real checkouts, the classifier that sorts ~500K merchants into 50 to 80 platform families, and the scheduler that decides which stores to test and when. Agents write much of the code; you own the design, the failure modes, and the verdict on what ships.
Coverage as your number. Machine-tested checkout today reaches 21% of non-Shopify merchants; you own the line from there past 80% across roughly 500,000 stores. This is the number an AI agent is really buying when it decides to trust us.
Fleet economics. Cost per verified checkout, held below the commission each check protects. You make the spend legible and make the fleet earn its keep, store by store, rather than making it small.
Anti-bot navigation. The evolving contest with fingerprinting, rate limits, and challenge walls — navigated at scale without breaking the store or the law.
The instrumentation that proves it. Dashboards and ledgers that show, for any claim, when it was last tested, whether the robot really reached the cart, and what the check cost. Correctness you can watch, not correctness you assert.
The number, co-signed. Within your first quarter you co-sign a seat charter — the model we run for senior operators. It names one machine-checkable number that proves the seat works (machine-tested checkout coverage is the obvious one) and writes down what you decide freely versus what you propose for the founder to sign. You own a number, not a backlog.

Who You Are
You reason in invariants, failure modes, and tradeoffs. Handed a checkout flow you have never seen, you can sketch the three ways it will break before you write a line. You see the platform family behind a one-off store, and the shared recipe behind a hundred one-off stores. When a robot fails at 2 a.m., your first question is structural: what class of store did we just discover?

You move fluidly between architecture and shipped code — a classifier design in the morning can be a deployed test by night — and you are as comfortable deciding what to build as how. You treat agents as leverage you verify, not autocomplete you trust: you can point at a system you shipped, name the hardest failure you personally diagnosed in it, and say what you changed. You can do this job by hand and prove it, and that mastery is exactly what lets you direct agents and trust — or reject — what comes back. The expensive thing here is a redo cycle, never the compute.

You have built browser automation, web scraping, or crawling systems at real scale, and you operated them in production — you know what a fleet of headless browsers does to your infrastructure bill and your on-call sleep. You have reverse-engineered a site that did not want to be automated, and won. Playwright, headless Chrome, proxy rotation, and queue-backed job systems are familiar ground; Node.js and Python are daily tools. We care about the artifact and the reasoning far more than where you did it — no degree to check, no pedigree to clear.

Who this isn't for. This is wrong if you guard a single lane and call the rest someone else's department — you own the fleet across automation, classification, infrastructure, and cost, and "that's not my job" ends the conversation. It's wrong if you pick technologies for how they'll read on your next resume rather than for what the fleet needs tonight. It's wrong if you wait to be told what to test instead of reading the system and deciding. And it's wrong if your code is whatever the model handed you and you couldn't say why it's right, or if you're comfortable letting an agent grade its own work. You'll be happiest here if your idea of craft is a fleet of robots that quietly proves, store after store, that a code is real.

Original source