Senior Manager, Systems Operations
- Location
- Atlanta, GA
- Posted
- 2026-08-03
Job description
Responsibilities
Team Leadership: Lead a globally distributed team of Analysts across the US, India, and Australia. Provide hands-on guidance, mentoring, and performance management. Hold a high bar for accountability and foster a culture of ownership, continuous improvement, and empowerment.
Process Improvement: Identify and drive process improvements that reduce incident frequency, eliminate operational toil, and increase team velocity. Champion automation adoption and data-informed decision-making as core to how the team operates.
Training and Support: Build team capability through structured training, clear documentation standards, and active coaching. Act as the primary escalation point for staff issues and drive a culture where knowledge is shared, not siloed.
Collaboration and Communication: Build strong working relationships with development, engineering, infrastructure, and network teams. Represent Systems Operations in cross-functional discussions and act as a credible operational partner to internal stakeholders.
Continuous Improvement: Lead or participate in continuous improvement initiatives that advance operational maturity. Actively identify and scope use cases for automation and AI-enabled workflows within the team’s day-to-day operations.
On-Call Rotation: Participate in an on-call rotation as needed.
Additional Duties: Perform any other activities as directed by management.
Knowledge and Experience
Experience: 5+ years of direct people management experience in a technical operations environment, with clear accountability for team performance, development, and retention. Team lead or delegation-only roles do not satisfy this requirement.
Strategic Thinking: Demonstrated ability to think beyond the ticket queue – connecting operational metrics to business outcomes, identifying systemic risks, and building plans that improve team capability over time.
Production Operations: Experience with managing production operations, monitoring, alerting, notifications, etc.
Scripting Languages: Proficiency in scripting languages such as Bash, Python, and/or PowerShell.
Server Administration: Strong proficiency with Linux and Windows Server administration.
Monitoring Tools: Experience with monitoring and alerting tools (Splunk, Nagios, BigPanda, PagerDuty).
Problem-Solving: Excellent problem-solving and troubleshooting skills.
Documentation: Process-oriented with great documentation skills (Confluence).
Business Continuity: Experience with automation of business continuity/disaster recovery.
Preferred Knowledge and Experience
Cloud Services: Experience with open-source technologies and cloud services (AWS/Azure).
Scheduling Tools: Experience with Rundeck and/or Cisco Tidal Enterprise Scheduler.
Monitoring Tools: Experience with BigPanda and PagerDuty.
AI Ops: Familiarity with AI-assisted operations – including workflow automation, AI-powered alerting and event correlation, and the use of AI tools to reduce operational toil. Demonstrated curiosity about applying AI to real operational problems is strongly valued.
Financial Markets / FIX Protocol: Experience in financial markets operations and/or FIX Protocol. Valuable for ramp speed and platform credibility – not a barrier to entry. Training and mentorship available on the platform side.