Senior Site Reliability Engineer, Production Engineer
- Company
- Cisco ThousandEyes
- Location
- Remote
- Work type
- Full Time
- Posted
- 2026-08-07
Job description
Responsibilities
Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
Participate in and improve our 24x7 incident response and on-call rotation.
Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
Automate production operations to provide guardrails and continuous platform operation.
Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.
Minimum Qualifications
5+ years of experience in a related role
Proficiency in software development with languages such as Python or Go
Shown ability to build and implement scalable, well-tested, and security-focused solutions that integrate security protocols throughout the development and deployment lifecycle
Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols
Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs
Preferred Qualifications
Familiarity with procedures for operating a large-scale, highly available enterprise platform
Excellent communication and documentation skills
Strong sense of ownership, drive, and attention to detail
Expert-level knowledge of Kubernetes and its ecosystem
In-depth knowledge of cloud providers, preferably AWS
Skills Required
5+ years of experience in a related role
Proficiency in software development with languages such as Python or Go
Ability to build and implement scalable, well-tested, security-focused solutions integrated across development and deployment lifecycle
Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols
Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs
Familiarity with operating large-scale, highly available enterprise platforms
Excellent communication and documentation skills
Expert-level knowledge of Kubernetes and its ecosystem
In-depth knowledge of cloud providers, preferably AWS