The role
From JPMorgan Chase's own posting.
Join JPMorganChase's Machine Learning Center of Excellence — where cutting-edge engineering meets real-world impact at global scale.
As a Lead Software Engineer at JPMorganChase within the Corporate Sector – Artificial Intelligence and Machine Learning Data Platforms and Machine Learning Center of Excellence team, you serve as a seasoned member of an agile team to design and deliver trusted, market-leading technology products in a secure, stable, and scalable way. You are responsible for carrying out critical technology solutions across multiple technical areas within various business functions in support of the firm's business objectives. You will collaborate with a multi-disciplinary community of experts focused exclusively on machine learning, working with cutting-edge techniques in disciplines such as deep learning and reinforcement learning.
Job responsibilities
Design, develop, and maintain production-grade Python services and APIs that power high-impact machine learning platforms at enterprise scale
Architect and implement high-throughput, low-latency distributed systems within Amazon Web Services (AWS) environments to support critical business workloads
Build and manage scalable cloud-native applications leveraging AWS technologies including Elastic Kubernetes Service (EKS), Elastic Container Service (ECS), Managed Streaming for Apache Kafka (MSK), Simple Queue Service (SQS), and S3
Develop reusable service frameworks, shared libraries, and modular application components that accelerate engineering delivery across teams
Design and implement infrastructure-as-code solutions using Terraform and CloudFormation to enable repeatable, auditable, and scalable deployments
Create and maintain monitoring, alerting, and observability solutions utilizing platforms such as Datadog, Dynatrace, and Splunk to ensure operational excellence
Deploy and support applications in production environments while ensuring adherence to service-level objectives and service-level agreements
Implement secure-by-design engineering practices, automated testing, and deployment strategies including blue/green and canary releases
Review code, provide architectural guidance, and mentor engineers on software engineering best practices to elevate team capability
Collaborate with product managers, platform engineering teams, and site reliability engineers to deliver scalable, business-aligned solutions
Drive adoption of enterprise-approved AI-assisted engineering practices to improve code quality, operational excellence, troubleshooting, and delivery efficiency
Required qualifications, capabilities, and skills
Formal training or certification on software engineering concepts and 5+ years applied experience
Advanced proficiency in Python programming, object-oriented design, and modular software architecture
Experience building and operating large-scale, high-performance cloud-native services within AWS environments
Hands-on experience with AWS technologies including EKS, ECS, MSK (Kafka), SQS, and S3
Strong experience implementing infrastructure-as-code solutions using Terraform and/or CloudFormation
Expertise in designing, deploying, and supporting distributed systems in production environments
Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Dynatrace, and Splunk
Strong understanding of API design, microservices architecture, and scalable system design patterns
Experience implementing automated testing, CI/CD pipelines, deployment automation, and secure software engineering practices
Demonstrated experience utilizing approved AI-assisted software development tools for coding, code review, testing acceleration, troubleshooting, and operational support
Strong understanding of responsible AI usage, application security, resiliency requirements, compliance standards, and mentoring engineers on engineering best practices
Preferred qualifications, capabilities, and skills
Strong knowledge of distributed systems reliability patterns, including resiliency engineering, self-healing architectures, backpressure management, and idempotency
Experience optimizing real-time and event-driven architectures at scale, particularly with Kafka-based messaging systems
Experience implementing end-to-end observability, automated operational runbooks, and proactive monitoring frameworks
Familiarity with CI/CD best practices, canary deployments, blue/green deployment strategies, and release automation within cloud environments
Familiarity with Generative AI and Large Language Model technologies and experience building engineering solutions that leverage AI/LLM platforms