Senior Site Reliability Engineer
- Location
- Fort Lauderdale, FL
- Work type
- Full Time · Hybrid
- Posted
- 2026-09-16
Job description
Job Description
Avaloq's R&D Lab is building a SaaS, API-first, composable banking platform. As a Senior Site Reliability Engineer you will help build the reliability practice behind it.
This is amongst the first technical roles we are hiring for our Fort Lauderdale site. You will partner with our Senior SRE in Zurich and work alongside our platform engineers and product teams.
The first three to six months:
Some of what you can expect early on:
Learning our AWS environment, our serverless platform, and how our product teams build and release
Contributing to the direction on observability tooling, together with the platform engineers, the SRE team, and our developers
Working with the Zurich team on the incident and on-call operating model
Taking on increasing responsibility as you build context, always in alignment with the Zurich team
What you will do:
Define and evolve SLIs, SLOs, and error budgets with product teams, and use them to drive reliability decisions
Collaborate on the design of our observability approach for a distributed serverless system, covering metrics, logs, and traces
Build the incident response practice with us: on-call model, escalation, blameless post-mortems, and the loop that turns findings into hardening work
Build reliability automation that detects and remediates issues before they reach clients
Improve CI/CD pipelines and deployment automation on GitHub Actions to reduce operational toil and release risk
Work with product teams on resilient design, capacity planning, and progressive delivery approaches such as canary and blue-green
Contribute to disaster recovery design and testing for client-facing environments
Partner with Security and Compliance so that operational practice holds up to regulatory and audit expectations
Help us define and implement operational readiness for client go-live
Mentor colleagues and support a culture of shared operational ownership
Our stack:
Serverless-first on AWS: Lambda, DynamoDB, SQS, S3, and Bedrock. Terraform and OpenTofu for infrastructure, GitHub Actions for CI/CD. Product teams work in Rust, TypeScript, Python, Vue, and Angular. We do not run containers.
Qualifications
5+ years in Site Reliability Engineering, DevOps, or production operations for distributed cloud systems, with substantial hands-on AWS experience
Practical experience defining and working with SLIs, SLOs, and error budgets
Strong incident response background, including on-call, triage under pressure, and post-mortem practice
A clear point of view on observability for distributed systems, and the reasoning behind it
Solid automation and scripting ability. We are language-agnostic; Rust, TypeScript, or Python all work here
Excellent collaboration and communication skills, with a pragmatic approach to balancing speed and stability
It would be a real bonus if you have:
Experience with AWS serverless and event-driven architecture, including idempotency, retries, dead-letter queues, and failure handling. Strong AWS generalists who want to go deep here are welcome
Experience taking a platform from pre-launch to production, including defining operational readiness criteria
Experience building AI agents to monitor platform health
Strong CI/CD background, ideally with GitHub Actions, including secrets management, least-privilege permissions, and deployment controls
Infrastructure as Code experience with Terraform or OpenTofu
Experience with progressive deployment approaches such as canary, blue-green, or rollback automation
Disaster recovery design and testing for production SaaS
Experience operating across more than one cloud provider, or designing for portability
Experience operating a B2B SaaS platform in a regulated environment, including audit-sensitive release processes
Familiarity with SOC 2, PCI DSS, or GDPR as they apply to operational practice
Experience with AI-assisted engineering tools
AWS certification, particularly DevOps Engineer - Professional, Solutions Architect - Professional, or Security - Specialty