← Back to jobs

Senior Voice AI Engineer

Company
Carbon Technology
Location
Remote
Work type
Full Time
Posted
2026-08-07

Job description

What you'll own
Voice, end to end. The Retell → LiveKit migration, then the platform every agent runs on — turn detection, interruption, SIP/WebRTC, multi-agent handoff, multi-tenant config. Migration succeeds when customers never notice. And it's not one receptionist: inbound, outbound, interviews, check-ins, surveys — a hive of voice agents, each pushing a different limit.

Past the APIs. Train and fine-tune our own voice models. Own inference and drive cost-per-conversation down as volume explodes. Ground agents in real customer data — Supabase, pgvector, retrieval into HiveMind, our intelligence layer that gets smarter with every call. Prove it with evals: "first-response accuracy went up 12% last week" — and show the data.

Wherever it takes you. Voice is the front door, not a cage. We keep the team tiny on purpose and automate everything else, so the people on it touch everything — backend, orchestration, product, whatever the mission needs. Find a better path than the one we planned? Take it. That's the job.

How we work
We're real coders. We were shipping long before AI wrote a line of code — but we embraced it early, and now one engineer here ships like five. If your instinct on a boring problem is "I'll build an agent so nobody ever does this again," you'll feel at home.

Ship fast. Remote-first, async-preferred, output over optics. No committees between you and production.

We run on EOS. One Team. Own It. Trajectory Wins. Reinvent What's Possible. Not wall art — how we decide.

Our stack
LiveKit Agents (Python) · Twilio + LiveKit SIP/WebRTC · multi-model LLMs (Anthropic, OpenAI, Google) · Deepgram / Cartesia / ElevenLabs today — our own voice models next · TypeScript + n8n orchestration · Next.js + Supabase (Postgres, pgvector, RLS) · HiveMind learning from every transcript

Who you are
You've shipped real-time voice AI to production and kept it alive at volume — LiveKit ideal, but if you built it on Retell, Vapi, Pipecat, or raw WebRTC pipelines, we care what you shipped, not which logo. Strong Python in streaming, async, real-time systems. You've built the full loop — STT, reasoning, TTS, VAD, turn-taking, tool calls — and debugged it at 2am.

Bonus points: eval frameworks for LLM systems, voice model fine-tuning, inference and cost optimization, multi-tenant SaaS, telephony scars.

Most of all: you read "that's not supported" as an invitation.

Structure: Contract-to-Hire
We're not asking you to quit your job and hope.

Phase 1 — 3-month paid contract, real production work from week one. Rate negotiated to fit your situation, benchmarked against the full-time band below. 20 hrs/week while transitioning is fine. Either side can end it — no hard feelings.

Phase 2 — Full-time conversion on mutual fit: $150–200K base + real equity as one of our earliest engineers, up to $20K annual bonus, health insurance, flexible PTO. Path to senior IC or technical leadership.

Exceptional candidate who needs to go straight to full-time? Let's talk, the structure serves the fit, not the other way around.

Deliberately funded. We raise carefully, and only what the mission needs, so the people building this keep real ownership and a real outcome. Long runway, no growth-at-all-costs pressure, and equity that's meant to matter.

Successful Founders with exits and ex-Apple leadership.

Skills Required
Experience shipping real-time voice AI systems to production and operating them at scale
Strong Python expertise in streaming, async, real-time systems
Built the full voice stack: STT, LLM reasoning integration, TTS, VAD, turn-taking, and tool calls
Experience with SIP/WebRTC and telephony platforms (Twilio) and LiveKit or equivalent real-time voice platforms
Experience integrating and operating LLMs and multi-model stacks (Anthropic, OpenAI, Google)
Experience with vector retrieval and databases (Postgres, Supabase, pgvector) and multi-tenant data concerns (RLS)
Experience training and fine-tuning voice models, inference optimization, and building eval frameworks for LLM systems
Familiarity with Deepgram, Cartesia, ElevenLabs or similar STT/TTS services
Experience with TypeScript, Next.js, and orchestration tooling (n8n) for full-stack or orchestration work
Startup experience and willingness to operate across backend, orchestration, product, and production support

Original source