Infrastructure Engineer
- Location
- New York City, NY
- Work type
- Full Time
- Posted
- 2026-08-03
Job description
🦸🏻♀️ What you’ll do
Technical
Breadth across disciplines. Bring deep focus to one problem at a time, with the breadth to move between SRE, DevOps, Infrastructure, and Platform work over a quarter or two as the leverage shifts. This is not a thrash-every-week role — most of the time you’re heads-down on one substantial initiative (the on-call posture, the release pipeline, the multi-region Terraform layout, the internal platform surface). Cross-layer fluency is what lets you pick the right next initiative; it isn’t a weekly context-switch.
Simplicity / via negativa. Challenge the status quo and remove toil before adding features — automate operational tasks and infrastructure management with Python or Go, reject tools that don’t fit the problem, and treat manual on-call work as a defect to be designed out, not a status quo to be staffed up.
Breadth across the stack. Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and the supporting cloud and AI tooling that backs WRITER’s high-traffic platform.
AI in workflow. Run agents in your daily loop — Claude Code, Droid, Codex, internal skills — to investigate incidents, draft Terraform / Helm changes, write runbooks, scaffold tooling, and review PRs. Build the agentic setup as a collective surface: humans and digital teammates working as one team, with shared skills, shared context, and shared on-call workflows. Encode recurring infra tasks as internal skills any teammate (human or agent) can pick up and run, so the team’s throughput compounds — not just your own.
Debugging fluency. Lead incident response, post-mortems, and root-cause analyses — trace failures to the underlying problem (never the symptom), apply the learning back into the architecture, and prevent the same incident from happening twice.
Non-technical
End-to-end ownership. Own the reliability, performance, and efficiency of WRITER’s core services end-to-end — define and uphold the SLOs and error budgets, carry the on-call pager, and stand behind the outcome metric, not just the system you shipped.
Strategic vs. tactical balance. Balance this week’s critical work with the 6–12-month platform direction — ship the on-call-driving fix today while shaping the multi-year observability, cost, and reliability investments that move WRITER’s enterprise customers.
Cross-functional collaboration. Operate at the seams with product, security, and engineering peers — provide expert guidance on system design for reliability, performance, and scalability from conception through launch, Connect the infra agenda to product and revenue context, and disagree with evidence, not volume.