The role
From Ally Bank's own posting.
Ally and Your Career
Ally Financial only succeeds when its people do - and that’s more than some cliché people put on job postings. We love this stuff! We see our people as, well, people - with interests, families, friends, dreams, and causes that are all important to them. Our focus is on the health and safety of our teammates as well as work-life balance and diversity and inclusion. From generous benefits to a variety of employee resource groups, we strive to build paths that encourage employees to stretch themselves professionally. We want to help you grow, develop, and learn new things. You’re constantly evolving, so shouldn’t your opportunities be, too?
Work Schedule : Ally designates roles as (1) fully on-site, (2) hybrid, or (3) fully remote. Hybrid roles are generally expected to be in the office a certain number of days per week as indicated by your manager. Your hiring manager will discuss this role's specific work requirements with you during the hiring process. All work requirements are subject to change at any time based on leader discretion and/or business need.


The Opportunity
We are seeking a Director to lead Site Reliability Engineering (SRE) and Production Operations. This senior leader is accountable for the operational health, stability, resilience, and availability of production applications and platforms supporting Ally’s Automotive and Insurance businesses. The role defines and executes the strategy for production operations and reliability engineering, driving continuous improvement and partnering across engineering, product, infrastructure, and business teams to deliver secure, stable, and scalable services.
At Ally, you get a startup feel, but experience the benefits of a company that’s worked out the kinks and is fulfilling its purpose. We’re always evolving and see that as a good thing. From owning our work to seeing its impact in the real world, our team is relentless in finding new ways technology can help make experiences better and help people. We are problem solvers, we value diverse thinking, we support one another, and we challenge ourselves to think bigger in the journey to deliver customer-obsessed tech solutions. To read more about what our tech team does, be sure to visit our tech blog at ally.tech


The Work Itself
Key Responsibilities
Lead SRE strategy and production operations for critical application platforms, ensuring availability, resiliency, recoverability, and performance targets are consistently achieved.
Own and evolve the operating model for production support, including incident, problem, and change risk management, as well as service restoration across the application portfolio.
Drive adoption of SRE practices, including service level indicators (SLIs), service level objectives (SLOs), error budgets, operational readiness, and automation-first engineering approaches.
Define the target-state SRE operating model and organization, including capacity planning, skill mix, and sourcing strategy (employees vs. contractors), to ensure sustainable 24x7 coverage aligned with business growth.
Establish and institutionalize best practices across SRE and application sustainment, creating consistent, scalable standards for reliability engineering and operational execution.
Establish and monitor operational health metrics, using data to identify systemic risks, improve reliability, reduce incident volume, and shorten recovery times.
Provide executive leadership during major incidents, ensuring rapid coordination, clear communication, timely escalation, and durable corrective actions.
Lead post-incident reviews and problem management efforts to resolve root causes, eliminate repeat issues, and strengthen operational discipline.
Partner with product, engineering, infrastructure, and architecture teams to embed reliability, operability, and supportability into design, delivery, and release processes.
Influence senior leaders across engineering, infrastructure, and business functions—including peer organizations and one level above—to align on reliability strategy, operating models, and investment priorities.
Lead the evolution of traditional application sustainment toward a modern SRE-led model, ensuring a balanced transition that enhances reliability without disrupting critical support responsibilities.
Drive automation and tooling investments that reduce manual effort, improve observability, streamline support processes, and increase engineering efficiency.
Evaluate and quantify the impact of AI-driven operations (AIOps) and automation accelerators, driving data-informed adoption to improve reliability, efficiency, and cost outcomes.
Define standards for monitoring, alerting, logging, capacity planning, and production readiness to strengthen proactive issue detection and service resilience.
Influence cloud and platform transformation efforts by clarifying operational ownership, improving support models, and aligning reliability practices with modern engineering patterns.
Establish a clear point of view on centralized versus distributed SRE models, shaping organizational design decisions that balance scale, accountability, and alignment with Agile delivery teams.
Build, lead, and develop high-performing teams, fostering accountability, technical depth, and a culture of continuous improvement.
Provide strategic guidance and technical assessments to senior leadership, translating operational risk and technology opportunities into clear, actionable business decisions.
Oversee vendor and partner relationships supporting production operations, ensuring service quality, accountability, and alignment with enterprise standards.
Champion operational excellence by challenging legacy practices, advancing reliability engineering maturity, and promoting modern support models.
Provide leadership accountability for 24x7 production operations, including direct engagement in major incidents and crisis events, demonstrating experience operating in high-availability, always-on environments.
What Success Looks Like
Success in this role is defined by stronger operational resilience, modernized reliability practices, and a high-performing organization that delivers consistent outcomes at scale. This leader will establish a proactive, engineering-led model for production operations that improves stability, reduces risk, and enables business growth.
Application availability, resiliency, and recovery consistently meet or exceed defined service targets for critical business services.
Incident volume, repeat issues, and time to restore service are reduced through stronger operational discipline and targeted engineering improvements.
SRE practices are embedded across the portfolio, with measurable adoption of SLOs, observability standards, automation, and production readiness.
A clearly defined and scalable SRE operating model is established, with the right balance of skills, capacity, and sourcing to support long-term needs.
Operational decisions are data-driven, supported by clear health metrics, risk indicators, and executive reporting.
Cross-functional teams demonstrate stronger accountability for reliability and improved alignment between delivery velocity and operational stability.
Manual effort and operational toil are reduced through automation, improved tooling, and streamlined processes.
Measurable gains in efficiency and effectiveness are achieved through adoption of AI-driven tooling and modern operational practices.
The organization demonstrates stronger readiness for growth, change, and platform modernization.
The function is recognized as a strategic partner that improves customer experience, protects business operations, and elevates enterprise reliability maturity.
Minimum Qualifications
9+ Years of Relevant Experience
Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent
Highly Preferred Qualifications
10+ years of experience in Site Reliability Engineering, Product