The role
From JPMorgan Chase's own posting.
As a Lead Software Engineer – Software Reliability at JPMorgan Chase as a part of our product team, you will design, build, and operate scalable, resilient systems using Python and modern reliability practices. You will apply Site Reliability Engineering (SRE) principles to improve availability, performance, and operational excellence, and you will help establish engineering standards that increase system reliability and security. You will collaborate with engineering and product partners to troubleshoot, optimize, and maintain production services while fostering a collaborative and inclusive team culture.
Job Responsibilities
Design and develop scalable and resilient systems using Python to support continuous improvement and apply Site Reliability Engineering (SRE) concepts to enhance system reliability and performance
Execute software solutions, including design, development, and technical troubleshooting and create secure, high-quality production code and maintain algorithms that run synchronously with appropriate systems
Produce or contribute to architecture and design artifacts, ensuring design constraints are met
Gather, analyze, and synthesize data to develop visualizations and reporting for software and system improvement
Identify hidden problems and patterns in data to drive improvements in coding hygiene and system architecture
Implement reliability engineering practices such as monitoring, alerting, and automated recovery
Define and measure Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to track system health
Conduct chaos engineering experiments to test system resiliency and identify weaknesses
Perform performance testing using tools such as JMeter to ensure scalability and stability
Collaborate with product teams to enhance system reliability, scalability, and performance and contribute to software engineering communities of practice and events exploring new and emerging technologies
Foster a team culture of diversity, opportunity, inclusion, and respect and participate in post-incident reviews and drive root cause analysis for system failures
Required qualifications, capabilities, and skills
Hands-on practical experience in system design, application development, testing, and operational stability and proficient in coding in Python
Experience developing, debugging, and maintaining code in a large corporate environment with modern programming languages and database querying languages
Knowledge of the Software Development Life Cycle, AWS cloud exposure, troubleshooting abilities, resiliency, and automation focus
Understanding of agile methodologies such as CI/CD, application resiliency, and security
Knowledge of software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, mobile, etc.)
Experience with AI and full understanding of the SDLC process, and MongoDB
Familiarity with reliability engineering concepts, including monitoring, alerting, and automated recovery and ability to implement and maintain system health checks and performance metrics
Commitment to writing maintainable, testable, and high-quality code
Understanding of SRE principles, including SLIs, SLOs, and error budgets and experience with incident response and root cause analysis
Experience with chaos engineering practices to test system resiliency and proficiency in performance testing tools such as JMeter
Preferred qualifications, capabilities, and skills
Familiarity with modern front-end technologies
Exposure to cloud technologies