← Back to jobs

AI Architect

Location
New York, NY
Work type
Full Time
Posted
2026-09-14

Job description

About the Role

We're looking for an AI Architect to lead the design, evolution, and scaling of our AI pipeline infrastructure. You'll take ownership of a backend platform built to orchestrate AI workflows at scale, and extend it to support a growing set of high-fan-out AI features that are central to our product strategy.

What You'll Inherit

Our AI infrastructure is built around an internal, API-only backend service responsible for orchestrating AI pipelines end to end. It runs on a modern web framework with a background job processing system, backed by a relational database, and is designed around durable, observable, idempotent jobs rather than ad hoc scripts. Some earlier AI workflows still run on a separate orchestration framework and are being progressively migrated into this platform.

Key infrastructure you'll own:

An internal, API-only backend application deployed on a cloud PaaS (staging and production environments)
A background job processing system with multiple queues and a durable message broker
A dedicated relational database for pipeline state and history
A job base class/pattern that provides idempotency guards, status-transition state machines, retry-with-backoff, structured logging, and error reporting on failure
Authenticated API access for inbound integrations, with secure credential management for outbound integrations
Multiple LLM providers for generation, structured output, and embeddings
A vector database for similarity search and matching at scale
Integrations with adjacent internal systems (content/CMS, notifications, messaging/chat tooling, tracing and error-monitoring platforms)

What you'll do

Own and Operate (Immediate)

Maintain and operate all existing AI pipelines running on the platform
Complete the migration of remaining workflows from the legacy orchestration framework to the primary platform
Own on-call response for AI pipeline failures, including failed-job triage and retries
Manage LLM provider relationships, API key rotation, cost tracking, and model upgrades

Architect and Scale (Ongoing)

Design and implement new high-fan-out AI pipelines on the existing platform — built to support future horizontal workflows (many generations across many users/items) without significant rework
Establish patterns and conventions for onboarding new pipelines, including registry entries, job subclassing, batch fan-out, and automated reporting
Drive architectural decisions around data storage strategy, queue partitioning, concurrency throttling, and cost controls as pipeline volume grows
Evaluate and integrate new LLM providers and embedding models as the landscape evolves, while maintaining backward compatibility with existing vector data
Build observability and operational tooling — extending the primary ops dashboard with custom reporting, cost tracking, and alerting as needed

Collaborate and Lead

Partner with product, editorial/content, and growth teams to translate product requirements into pipeline designs
Work across the engineering organization to integrate AI pipelines with the broader technical stack
Mentor engineers on AI pipeline patterns, prompt engineering, and platform architecture
Own the technical proposal process for new pipelines and major infrastructure changes

Required Qualifications

Deep backend framework expertise — 5+ years building production applications in a modern web framework (e.g., Rails, Django, or similar), with strong experience in ORM usage, API-only application design, and background job processing
Job orchestration at scale — Proven experience designing idempotent, retryable, fan-out job pipelines (batching, concurrency controls, dead-letter handling, queue partitioning)
LLM integration experience — Hands-on work with multiple LLM providers, structured output, embeddings, prompt engineering, and cost tracking
Vector search infrastructure — Experience with a vector database for similarity matching at scale (batched queries, namespace management, embedding model migration)
Production operations mindset — Experience running internal services on a cloud PaaS or equivalent, with relational databases, caching/queueing infrastructure, and observability tooling
Architecture and proposal-driven development — Track record of authoring technical design documents, making build-vs-buy decisions, and designing platforms meant to be extended by others

Preferred Qualifications

5+ Years of Ruby/Ruby on Rails experience
Experience migrating workflows from a Python-based orchestration framework to Ruby-based framework
Familiarity with agent orchestration and LLM tracing/observability ecosystems
Experience with AI alerting or notification systems (change detection, embedding-based matching, threshold tuning)
Background in media/publishing AI applications (personalization, content recommendation, cohort-based delivery)
Experience with state machine libraries for managing job lifecycles
Experience building retrieval-augmented generation (RAG) pipelines - grounding LLM outputs in retrieved documents or context, not just retrieval-based matching/ranking

What Success Looks Like

First 90 days: All existing pipelines are stable and well-understood. Remaining migrations are scoped and underway. You've shipped at least one new pipeline or major platform improvement.

First year: The platform is the single home for all AI orchestration. New pipelines are onboarded in days, not sprints. LLM costs are tracked and optimized. The team trusts the platform's reliability and observability.

Original source