← Back to jobs

Senior Site Reliability Engineer

Company
PENN Entertainment, Inc.
Location
United States
Work type
Full Time · Remote
Posted
2026-08-30

Job description

About the Role & Team
The SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical infrastructure across compute, networking, storage, and cloud services (GCP/AWS) — driving complex migrations, building platform tooling and automation (ArgoCD, Helm, GitHub Actions), and improving observability and incident response across hundreds of production services spanning multiple regulated jurisdictions. We're looking for someone with strong Kubernetes and distributed systems experience, proficiency in Go, Python, or Bash, and a track record of leading cross-team infrastructure projects with high autonomy. You'll solve ambiguous problems, reduce operational toil through automation, mentor teammates, and bring a production-first perspective to architecture decisions — all on a team that values pragmatic engineering, real ownership, and continuous improvement of how we work.

About the Work

Drive complex infrastructure migrations and projects — scoping, planning, execution, and validation across multiple production environments and jurisdictions
Build and maintain platform tooling and automation — ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows that reduce toil for SRE and development teams
Support development teams — consult on infrastructure needs, unblock cross-team dependencies, review architecture proposals, and help teams adopt platform tooling and best practices
Design and improve observability and alerting — Datadog monitors, dashboards, and runbooks that surface meaningful signals and make systems operable by the whole team
Provide operational support and incident response — investigate and resolve production issues through structured debugging and root cause analysis, and contribute to on-call rotations to maintain platform reliability

About You

5+ Years of Experience in a similar role (DevOps, Site Relatability Engineer)
Strong experience operating and troubleshooting Kubernetes in a production Linux environment (cluster lifecycle, networking, storage, scheduling)
Experience working with AWS, GCP, and/or on-premise environments
Proficiency in at least two of: Go, Python, Bash/Shell — for building tooling, automation, and debugging production systems
Deep understanding of distributed systems — failure modes, networking fundamentals, capacity planning, and performance analysis
Experience with GitOps and CI/CD workflows (ArgoCD, Helm, GitHub Actions, or similar)
Experience with infrastructure-as-code (Terraform, Helm, or equivalent)
Track record of leading complex migrations or infrastructure projects with cross-team dependencies
Strong incident response and troubleshooting skills — structured debugging across multiple services and infrastructure layers
Clear technical communication — you write docs others can use, explain trade-offs to varied audiences, and proactively unblock cross-team dependencies

Nice to have
Experience with service mesh technologies (Istio, Cilium)
Familiarity with distributed storage systems (Ceph, or similar)
Experience with bare-metal Kubernetes or Talos OS
Exposure to regulated environments (sports betting, fintech, or similar compliance-heavy domains)
Experience with Datadog or comparable observability platforms at scale
Familiarity with database operations — PostgreSQL, connection pooling (PgBouncer), or database migration tooling

Original source