Site Reliability Engineer
- Location
- New York, NY
- Work type
- Full Time · On-site
- Posted
- 2026-09-03
Job description
The Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.
We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate software in public cloud environments. Our cross-functional teams span the stack, from front ends to APIs to databases, and they have all the skills and resources they need to build, ship and operate their own software.
Skills
Devops, Cloud, Aws, Terraform, Automation, incident management
Top Skills Details
Devops ,Cloud, Aws, Terraform, Automation ,incident management
Additional Skills & Qualifications
Some of the problems we'll work on include:
- Supporting application in production, including incident response and post-incident reviews
- Applying observability engineering to our applications to ensure we can proactively detect system degradation, easily understand system state, and quickly diagnose issues
- Investigate and resolve production issues
- Building automation to reduce toil and improve developer productivity
As a Site Reliability Engineer in our group, you will:
- Scope technical projects and break them down into user stories and tasks within an engineering team
- Directly contribute to the design and coding of our software systems.
- Contribute to Build systems that are secure, reliable, scalable, and extensible
- Make sound technical decisions utilizing the advice of teammates and contribute to technical conversations with other engineering teams
- Build and maintain CI/CD pipelines to automate the deployment of our software
- Automate the provisioning and management of our infrastructure using Infrastructure as Code (IaC) tools
- Define, implement, and maintain observability solutions for our applications to ensure we can proactively detect system degradation, easily understand system state, and quickly diagnose issues
- Diagnose and resolve production issues, including performance tuning and capacity planning