← Back to jobs

Staff , Site Reliability Engineer – Cloud Platform

Location
New York City, NY
Work type
Full Time
Posted
2026-08-03

Job description

What you’ll do
Own observability end to end. Evolve and streamline our metrics, logs, and distributed-tracing strategy. Building telemetry standards that give every team a clear, real-time picture of production health
Drive reliability through SLOs. Establish meaningful service-level objectives with product and engineering teams, and use error budgets to balance velocity against stability.
Lead incident response. Help set the standards for our incident response feedback loop. Manage the on-call rotations for production systems, respond to incidents with calm and rigor, manage blameless postmortems – then systematically reduce toil.
Engineer for reliability. Operate and improve our Kubernetes/EKS workloads on AWS. Build automation and tooling that removes manual operational work and improves developer productivity.
Be a force multiplier. Mentor engineers on operational excellence, set standards for how services are built to be observable and reliable, and influence architecture decisions across teams.

Qualifications
Baseline skills/experiences/attributes:

8+ years of managing production systems
Deep, hands-on production experience with AWS and Kubernetes
Strong programming/scripting skills and experience building operational automation
Proven ownership of a full stack observability platform (NewRelic/Datadog/etc)
A track record of leading incident response and delivering improvements in reliability
Clear written and verbal communication skills
Strong desire for being a mentor to others

Ideally, you also have these skills/experiences/attributes (but it’s ok if you don’t!):

Background operating in a highly regulated or compliance-sensitive environment (e.g., HIPAA, SOC 2, FedRAMP, SaMD, etc).
Background in technical Healthcare Protocols (DICOM, HL7, FHIR) and systems integrations (PACS, VNA, EMR, etc)

Original source