Senior Database Reliability Engineer
- Company
- PLACE
- Location
- United States
- Work type
- Full Time · Remote
- Posted
- 2026-09-15
Job description
What You're Great At:
You're the rare engineer who is equal parts careful and ambitious. You think about what can go wrong before it does and always have a mitigation ready — but you're not content to just hold the line. You see what a platform could become and build toward it. You've spent 10+ years in production database engineering, including real operational ownership of a 24x7 environment, and you've personally executed and measured restores under pressure rather than just watched backup jobs succeed. You know SQL Server 2016+ inside and out — Always On availability groups, failover clustering, transactional replication — and you're just as comfortable in advanced T-SQL, execution plans, and indexing strategy, with the judgment to know when the real fix lives in the schema rather than the query. You've run SQL Server in AWS (RDS, RDS Custom, or EC2) and have at least one on-prem-to-cloud migration under your belt. You've managed SQL Agent jobs, automated with PowerShell and/or Python, and kept legacy SSIS packages and SSRS reports dependable. You've worked at multi-terabyte scale, and you communicate proactively — stakeholders hear about degradations from you before they have to ask.
What You'll Do:
Own availability, recoverability, and performance for all SQL Server platforms in the acquired data estate
Build and maintain a complete, documented inventory of the SQL Server estate — instances, editions, licensing, sizes, growth rates, dependencies, job schedules, and consumers
Validate the backup strategy against agreed Recovery Point Objectives; perform and report scheduled test restores that prove Recovery Time Objectives
Manage high availability (Always On, failover clustering, log shipping, transactional replication), including planned failover exercises on a published cadence
Ratify service level objectives per system tier based on measured baselines, not aspiration
Lead incident response for database-layer events — triage, mitigation, communication, and written post-incident reviews with tracked corrective actions
Own disaster recovery and contingency planning, including annual DR exercises
Administer, troubleshoot, and extend legacy SSIS packages and SSRS reports; develop new SSIS, T-SQL, and stored procedures where needed
Automate provisioning, patching, index/statistics maintenance, integrity checks, backup verification, and environment refreshes via PowerShell, T-SQL, Python, or IaC
Build platform telemetry and alerting covering data quality, job durations, wait stats, blocking/deadlocks, storage growth, and cost
Write runbooks thorough enough that another operator can run the environment in your absence
Own database-layer access reviews, encryption, key rotation, credential management, and patching currency
Act as the go-to technical authority on the SQL Server estate; participate in design reviews and flag risk early
Mentor engineers and analysts, and cross-train at least one colleague to keep the on-call rotation sustainable
Participate in a shared on-call rotation supporting a 24x7 production environment (rotation composition, on-call compensation, and maintenance windows to be confirmed)
Preferred:
Experience migrating or retiring legacy Microsoft BI assets (SSIS, SSRS) into a modern cloud data platform
Familiarity with AWS DMS for replication and migration
Exposure to Snowflake, dbt, or comparable modern warehouse tooling
MySQL, PostgreSQL, or Aurora experience
Azure SQL experience
Experience operating through a post-acquisition integration
Property technology, mortgage technology, or fintech domain background